arXiv:2608.28458 · 工程化
Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
获取、修复、保留:面向小模型对话游戏 Agent 的诊断驱动后训练方案
Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
- 类型:arxiv
- 标识:2608.28458
- 链接:https://arxiv.org/abs/2608.28458
- 主分类:engineering
- 形态:benchmark
- TLDR:Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Challenge with a 2B open-weight model, and find that many failures are not only broad knowledge failures but also local decision failures: repeated guesses, malformed actions, and violations of feedback that the model has just seen. These diagnostics motivate a training recipe organized around three steps: acquire broad game participation through supe
- 副分类:agent
- 待LLM分类:否
- 标题中文:获取、修复、保留:面向小模型对话游戏 Agent 的诊断驱动后训练方案
- TLDR中文:交互式对话游戏检验了一种静态基准大多未明确覆盖的能力:模型需跨轮次维护状态、理解反馈,并在动态约束下选择合法动作。我们在 LM Playschool Challenge 上以 2B 开源权重模型研究该场景,发现许多失败不仅是广义知识欠缺,也包括局部决策失误:反复猜测、动作格式错误,以及违背刚刚收到的反馈。基于这些诊断,我们提出围绕三步组织的训练方案:通过监督 supe
- 来源文件:
- /inbox/tom/_candidates/2026-09-01-rag-retrieval-reranking-candidates.json
- /inbox/tom/_candidates/2026-09-01-agent-rag-longcontext-candidates.json