Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
- 类型:arxiv
- 标识:2609.07139
- 链接:https://arxiv.org/abs/2609.07139
- 主分类:engineering
- 形态:method
- TLDR:A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute the model must infer gradually over a conversation, namely how expert its dialogue partner is. Using ExpertCollab, a corpus of multi-turn research-planning dialogues between model-played personas at four expertise levels, we find that partner expertise is most decodable in the ear
- 待LLM分类:是
- 标题中文:Encoded Early, Used Late:Transformer 在何处开始作用于所推断的对话伙伴专业度
- TLDR中文:Transformer 可在残差流的某一深度使某个属性线性可解码,而此时该属性尚未影响输出。这种"可读位置"与"被使用位置"之间的差距已在输入中直接陈述的属性上得到验证。我们进一步追问:对于模型须在多轮对话中逐步推断的属性——即对话伙伴的专业度——这一现象是否依然成立。利用 ExpertCollab(一个由模型扮演的四级专业度角色进行多轮研究规划对话的语料库),我们发现伙伴专业度在……
- 来源文件:
- /inbox/tom/_candidates/2026-09-09-agent-rag-longcontext-candidates.json