📋 Tom 文献雷达 · Agent · RAG · 长上下文 · 2026-10-11
主题: Agent / RAG / 检索 / 长上下文 / 评测
来源: arXiv metadata + HF Daily(2026-10-07 ~ 10-08)
检索策略: all:"retrieval augmented generation" + HF 策展标签过滤
去重参考: 近 7 天文件名(2026-10-04 ~ 10-10)确认本期 8 条均为新条目
🔬 高价值条目(3 条)
1. Opera: A Verbal Critic Framework for Long-horizon Coding Agents
- 来源: HF Daily · arXiv 2609.33987(2026-10-07)· votes: 10
- 核心观点: 长周期 coding agent 需要及时修正,但现有 critic 往往在反馈发出后不再跟踪效果。Opera 将每次修正视为"持久笔记",持续跟踪直到问题真正解决,包含周期性 + 事件驱动触发、typed operator 诊断、交付前证据审计等机制。
- 关联标签:
agentbenchmark - 适用场景: coding agent 轨迹评估、反馈机制设计
2. ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI
- 来源: arXiv 2610.12415(2026-10-08)
- 核心观点: 用 RAG + 结构化 prompt 工程生成 malware 特定欺骗 playbooks,线下构建并验证后才部署,而非依赖运行时实时检测。知识库涵盖 malware 程序知识与欺骗编排逻辑。
- 关联标签:
ragbenchmark - 亮点: RAG 在网络安全对抗场景的创新应用,RAG 系统本身即为核心技术栈
3. Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
- 来源: HF Daily · arXiv 2610.11169(2026-10-07)
- 核心观点: Agent skills(SKILL.md)通过复制在 GitHub 仓库间传播,形成无注册表、无版本控制的软件供应链。首次构建了基于 git 历史的 agent skill 有向复制网络,可追溯来源、传播范围和安全修复覆盖范围。
- 关联标签:
agent - 亮点: 第一个系统梳理 AI agent skill 供应链的研究,对安全审计和依赖管理有直接价值
📡 其他候选(5 条)
4. The Geometry of Hierarchical Navigation: Accuracy and Query Cost for Point Process Input
- arXiv 2610.12312(2026-10-08)·
ragsystems - 研究 HNSW/层次图结构在高维向量空间贪婪导航的几何条件,确定性 coverage condition 保证任意查询收敛。对 RAG 召回系统有理论参考价值。
5. Is Memorization Context-Sensitive? Prefix-Based Extraction Beyond Isolated Prefixes
- arXiv 2610.12085(2026-10-08)·
ragbenchmarksystems - 在 RAG 实际部署中,前缀提取的记忆化风险是否被上下文条件缓解?通过配对 item 级测量证明上下文条件改变了可提取样本集合,但并未消除风险。对 RAG 安全边界有实测意义。
6. Forms of LLM-Integrated Applications: From LLM-Chats to Autonomous AI Agent System
- arXiv 2610.11899(2026-10-08)·
agentragsystems - 系统梳理 LLM 在软件系统中的集成形态:chatbot / copilot / RAG / workflow / coding agent / AI agent,揭示 vendor 标签背后的真实架构含义,copilot = router-worker,agent = AI 规划多步骤执行。
7. Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
- arXiv 2609.35855(2026-09-24)·
benchmarksystems - 将被拒绝的候选配置转化为后续优化的 stepping stones,而非丢弃,避免重复陷入相同失败模式。对 prompt/harness/代码自动优化流程有启发。
8. Hebero: GPU-Parallel Heterogeneous Multi-Task RL Benchmark
- HF Daily · arXiv 2606.03335(2026-10-05)· votes: 8 ·
benchmarksystems - GPU 并行 Isaac Lab 基准,40 个异构机械臂任务联合训练评测。DGPO 方法处理稀疏奖励与有限演示。偏向机器人,agent 评测可参考其多任务评估范式。
📦 候选 JSON
/shared/research-kb/inbox/tom/_candidates/2026-10-11-agent-rag-longcontext-candidates.json
ℹ️ 备注
- Substack:本轮搜索未发现匹配的近期高质量 Substack 研究帖(搜索词:agent RAG long context 2026),Cobus Greyling 帖子为旧文,本次未纳入。
- CSDN:本次无适用场景(未涉及版本/命令/源码/复现经验内容),未调用。
- 锁已持有,TTL 1500s,任务正常结束。