📡 Tom 文献雷达 · Agent × RAG × Long Context
日期: 2026-09-19 (20:40 UTC+8) | 轮次: 晚间档 | 来源: HF Daily (arXiv 富化失败)
💡 本期高价值(2 条)
1. ActObs — 预测环境观测让 RL 探索更优
- 来源: arXiv (2609.20715) | Votes: 15
- 标签:
agentRLmultimodal - 核心: 标准 SFT 只监督 action token,环境观测仅作上下文。ActObs 进一步监督轨迹中的 observation token,让策略在学习过程中建模动作后果,无需额外数据、参数或额外 forward pass。
- 意义: 对多模态 Agent 的环境感知和探索策略训练有直接参考价值。补充了"不要把环境当黑箱"的训练视角。
- 链接: https://arxiv.org/abs/2609.20715
2. Self-Evolving Search Index — RAG 索引自主适应环境
- 来源: arXiv (2609.19656) | Votes: 22 | (今日第三次上榜,持续高热)
- 标签:
agentragbenchmark - 核心: 检索质量高度依赖 index keys 对文档知识的暴露程度,有效表达随环境变化而不同。让索引自主诊断失败并迭代优化策略,减少人类干预,实现 RAG 索引层的自动化闭环。
- 意义: 对构建动态、自适应的 Agent 知识库有直接工程参考价值。
- 链接: https://arxiv.org/abs/2609.19656
📋 其余候选(6 条)
| # | 标题 | 票数 | 标签 |
|---|---|---|---|
| 3 | Verifiable Social Reasoning / Fuse | 29 | agent, benchmark |
| 4 | Can MiniMax-H3 Reason About the Physical World? | 91 | benchmark, multimodal |
| 5 | Sample Count Is Not Enough: Candidate-Generation Strategy | 14 | systems |
| 6 | VākQA: Telugu Spoken Factoid QA | 12 | benchmark |
| 7 | FAMOS: 3D Articulation from Sparse Observations | 17 | research |
| 8 | Srijika: OpenType Font Restyling 9 Indic Scripts | 14 | systems |
📌 本期说明
- arXiv API 今日返回 406,候选全部来自 HF Daily 富化流。
- Fuse 今日已两次上榜(0840、1440),高票(29)且评测价值高,保留在候选表。
- MiniMax-H3 票数最高(91),多模态物理推理评测,与 Agent/RAG 相关性偏弱但 benchmark 价值值得关注。
- Substack:本轮未主动检索。
Tom 文献雷达 · 每日 3 频 · 轻量版