DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users

  • 类型:arxiv
  • 标识:2610.03020
  • 链接:http://arxiv.org/abs/2610.03020v1
  • 主分类:agent
  • 形态:benchmark
  • TLDR:Long-term agents must remember not only what is true about a user, but also how a particular agent should work with that user as their shared history evolves. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving this relationship-specific agent memory implicit. Additionally, most prior works measure the model solely with final-answer QA over long interaction histories, making the assessment still incomplete and unreliable. To this end, we introduce DyadMem with the proposed new definition User-conditioned Relational Agent Memory (URAM)
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:DyadMem:关于 Agent 如何与用户协作的长期记忆基准
  • TLDR中文:长期 Agent 不仅需要记住关于用户的真实信息,还需要记住随着共同历史演进,特定 Agent 应如何与该用户协作。现有基准主要监督用户事实和偏好,或监督跨用户可复用的经验,使这种关系特定的 Agent 记忆未被显式建模。此外,大多数先前工作仅通过长交互历史上的最终答案问答来评估模型,导致评估仍不完整且不可靠。为此,我们提出 DyadMem 及新定义——用户条件关系型 Agent 记忆(URAM)。
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-05-agent-memory-tool-use-candidates.json