Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents

  • 类型:arxiv
  • 标识:2608.28458
  • 链接:https://arxiv.org/abs/2608.28458
  • 主分类:engineering
  • 形态:benchmark
  • TLDR:Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Challenge with a 2B open-weight model, and find that many failures are not only broad knowledge failures but also local decision failures: repeated guesses, malformed actions, and violations of feedback that the model has just seen. These diagnostics motivate a training recipe organized around three steps: acquire broad game participation through supe
  • 副分类:agent
  • 待LLM分类:否
  • 标题中文:获取、修复、保留:面向小模型对话游戏 Agent 的诊断驱动后训练方案
  • TLDR中文:交互式对话游戏检验了一种静态基准大多未明确覆盖的能力:模型需跨轮次维护状态、理解反馈,并在动态约束下选择合法动作。我们在 LM Playschool Challenge 上以 2B 开源权重模型研究该场景,发现许多失败不仅是广义知识欠缺,也包括局部决策失误:反复猜测、动作格式错误,以及违背刚刚收到的反馈。基于这些诊断,我们提出围绕三步组织的训练方案:通过监督 supe
  • 来源文件
  • /inbox/tom/_candidates/2026-09-01-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-09-01-agent-rag-longcontext-candidates.json