When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety

  • 类型:arxiv
  • 标识:2609.34771
  • 链接:https://arxiv.org/abs/2609.34771
  • 主分类:risk
  • 形态:method
  • TLDR:Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignm
  • 副分类:evaluation
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json