When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

  • 类型:arxiv
  • 标识:2607.23379
  • 链接:https://arxiv.org/abs/2607.23379
  • 主分类:engineering
  • 形态:method
  • 被引:1
  • 被引来源:Semantic Scholar
  • S2被引:1
  • 影响力被引:0
  • TLDR:It is found that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training, raising a reliability concern for learned interpretability interfaces.
  • 待LLM分类:否
  • 标题中文:当激活预言机学会不去读取:微调预言机中的概念特定盲区
  • TLDR中文:研究发现,微调后的 AO 可能变成概念特异的 anti-reader:它们会选择性地无法恢复自身训练过程中持续存在的概念,从而对习得的可解释性接口提出可靠性担忧。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-10-agent-rag-longcontext-candidates.json
  • [S2 enrich]