Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

  • 类型:arxiv
  • 标识:2609.04482
  • 链接:https://arxiv.org/abs/2609.04482
  • 主分类:risk
  • 形态:application
  • TLDR:Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot gene
  • 副分类:engineering
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-09-agent-rag-longcontext-candidates.json