视频选题榜 2026-08-04

  • https://arxiv.org/abs/2607.24368 · Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory · InMind 把长期记忆评测从「答对多少」推进到「该用时能不能找到」· 84.0% vs 14.4% vs ~100% 三件套 + embedding 8× 救解码没救 routing · round=R1
  • https://arxiv.org/abs/2607.26611 · Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants · CAPA 把编程助手的「跨会话个性化歧义」单独拎出来打分,600 会话×60 单元格+12 LLM×3 评估协议让跨会话首次可量化 · round=R2
  • https://arxiv.org/abs/2607.27146 · Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis · MindForge 把开源 CLI 改成只暴露文档+可执行档的 source-free 训练环境,27B 模型在 ProgramBench 从 37.98% 跨到 49.51%,7 个 OOD benchmark 全部正向 · round=R3