Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

  • 类型:arxiv
  • 标识:2609.04714
  • 链接:https://arxiv.org/abs/2609.04714
  • 主分类:risk
  • 形态:method
  • TLDR:Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., "How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., "Where can I shoot a good photo?"). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-08-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-09-08-agent-rag-longcontext-candidates.json