PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
- 类型:arxiv
- 标识:2607.05910
- 链接:https://arxiv.org/abs/2607.05910
- 主分类:risk
- 形态:benchmark
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:This work proposes PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt), and confirms that matched pass/block boundary pairs are essential for stable policy adaptation.
- OpenAlex ID:W7167679369
- OpenAlex DOI:10.48550/arxiv.2607.05910
- DOI:10.48550/arxiv.2607.05910
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2607.05910
- OpenAlex更新:2026-07-19
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
- TLDR中文:论文提出 PolicyShiftGuard,一个紧凑的策略条件护栏,采用结合随机策略 SFT(RP-SFT)与边界对策略适配(BP-Adapt)的两阶段训练方案,并验证匹配的通过/拒绝边界对是稳定策略适配的关键。
- 来源文件:
- /inbox/tom/_candidates/2026-07-16-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]