Tom 文献雷达 · Agent / RAG / Long-Context · 2026-09-02 14:40 UTC
本期候选(8条)
🔴 高价值(4条)
1. Safin-1: Safety from Within through Memory-Native State Evolution - 链接: http://arxiv.org/abs/2609.00092v1 - 来源: arXiv | 2026-08-31 | votes: 13 - 摘要: 提出 "Safety from Within" 范式,通过 Memory-Anchor Routing 让安全能力内嵌于模型自身计算,而非依赖外部对齐微调。适用于长时交互、状态累积场景。 - 标签: memory
2. Agents in the Large: Perception-Centered Architecture for Persistent Agents - 链接: http://arxiv.org/abs/2608.30478v1 - 来源: arXiv | 2026-08-31 | votes: 2 - 摘要: 研究持久化 AI Agent 框架,系统梳理记忆/工具/决策在长生命周期中的组织方式,提出感知中心架构,值得作为架构参考。 - 标签: agent, memory, systems
3. UI-Venus-2 Technical Report - 链接: https://arxiv.org/abs/2609.00028 - 来源: HF Daily | 2026-08-26 | votes: 47 - 摘要: 跨移动/Web/桌面统一 GUI Agent,170+ 多语言环境覆盖,统一闭环推理-动作框架。Bridge benchmark-to-production gap 的工程实践参考。 - 标签: agent, rag, benchmark, multimodal
4. Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement - 链接: https://arxiv.org/abs/2609.01481 - 来源: HF Daily | 2026-08-31 | votes: 3 - 摘要: 编码 Agent 在多天自主开发中持续自我改进的框架,组织为"测试-修复-增长"迭代循环,对 RAG + Agent 架构中的长期状态管理有参考价值。 - 标签: agent, systems
🟡 补充(4条)
5. SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers - 链接: https://arxiv.org/abs/2609.01343 | HF Daily | 2026-08-31 | votes: 48 - Scaling Laws + MoE + Looped Transformer,per-token FLOPs 匹配对比,是架构效率的工程参考。
6. DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory - 链接: https://arxiv.org/abs/2609.00768 | HF Daily | 2026-08-31 | votes: 8 - 自进化框架,用层级错误记忆指导问题生成方向,类比 RAG 反馈驱动的记忆积累。
7. ZimaBlue: Scalable Video Pre-training for World Action Models - 链接: https://arxiv.org/abs/2609.00188 | HF Daily | 2026-08-30 | votes: 32 - 视频预训练学习通用世界动作模型,egocentric video → robot control,embodied agent 方向。
8. Qwen-Drive-1.0: Vision-Language Foundation Model for Autonomous Driving - 链接: https://arxiv.org/abs/2609.00111 | HF Daily | 2026-08-30 | votes: 28 - VLM + 3D BEV + 规划统一框架,多模态 agent 感知规划参考。
趋势观察
- 记忆/状态演进 方向活跃:Safin-1(安全内嵌)、DiagEvo(错误记忆)、Agents in the Large(持久化)均指向同一核心问题——Agent 如何在长生命周期内累积、更新、检索记忆。
- GUI Agent 工程落地 加速:UI-Venus-2 覆盖 170+ 环境,是 benchmark → production 的重要实践。
- 编码 Agent 持续自我改进:Harness-of-Harness 代表了将测试-修复闭环内置于 Agent 的新方向。
附注
- 本期 Substack 检索无显著相关结果(web search 返回内容偏泛),未引入额外来源。
- CSDN 未使用。
- Candidates JSON:
/shared/research-kb/inbox/tom/_candidates/2026-09-02-agent-rag-longcontext-candidates.json