📋 Tom 文献雷达 · Agent · RAG · 长上下文 · 2026-10-11

主题: Agent / RAG / 检索 / 长上下文 / 评测
来源: arXiv metadata + HF Daily(2026-10-07 ~ 10-08)
检索策略: all:"retrieval augmented generation" + HF 策展标签过滤
去重参考: 近 7 天文件名(2026-10-04 ~ 10-10)确认本期 8 条均为新条目


🔬 高价值条目(3 条)

1. Opera: A Verbal Critic Framework for Long-horizon Coding Agents

  • 来源: HF Daily · arXiv 2609.33987(2026-10-07)· votes: 10
  • 核心观点: 长周期 coding agent 需要及时修正,但现有 critic 往往在反馈发出后不再跟踪效果。Opera 将每次修正视为"持久笔记",持续跟踪直到问题真正解决,包含周期性 + 事件驱动触发、typed operator 诊断、交付前证据审计等机制。
  • 关联标签: agent benchmark
  • 适用场景: coding agent 轨迹评估、反馈机制设计

2. ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI

  • 来源: arXiv 2610.12415(2026-10-08)
  • 核心观点: 用 RAG + 结构化 prompt 工程生成 malware 特定欺骗 playbooks,线下构建并验证后才部署,而非依赖运行时实时检测。知识库涵盖 malware 程序知识与欺骗编排逻辑。
  • 关联标签: rag benchmark
  • 亮点: RAG 在网络安全对抗场景的创新应用,RAG 系统本身即为核心技术栈

3. Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub

  • 来源: HF Daily · arXiv 2610.11169(2026-10-07)
  • 核心观点: Agent skills(SKILL.md)通过复制在 GitHub 仓库间传播,形成无注册表、无版本控制的软件供应链。首次构建了基于 git 历史的 agent skill 有向复制网络,可追溯来源、传播范围和安全修复覆盖范围。
  • 关联标签: agent
  • 亮点: 第一个系统梳理 AI agent skill 供应链的研究,对安全审计和依赖管理有直接价值

📡 其他候选(5 条)

4. The Geometry of Hierarchical Navigation: Accuracy and Query Cost for Point Process Input

  • arXiv 2610.12312(2026-10-08)· rag systems
  • 研究 HNSW/层次图结构在高维向量空间贪婪导航的几何条件,确定性 coverage condition 保证任意查询收敛。对 RAG 召回系统有理论参考价值。

5. Is Memorization Context-Sensitive? Prefix-Based Extraction Beyond Isolated Prefixes

  • arXiv 2610.12085(2026-10-08)· rag benchmark systems
  • 在 RAG 实际部署中,前缀提取的记忆化风险是否被上下文条件缓解?通过配对 item 级测量证明上下文条件改变了可提取样本集合,但并未消除风险。对 RAG 安全边界有实测意义。

6. Forms of LLM-Integrated Applications: From LLM-Chats to Autonomous AI Agent System

  • arXiv 2610.11899(2026-10-08)· agent rag systems
  • 系统梳理 LLM 在软件系统中的集成形态:chatbot / copilot / RAG / workflow / coding agent / AI agent,揭示 vendor 标签背后的真实架构含义,copilot = router-worker,agent = AI 规划多步骤执行。

7. Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution

  • arXiv 2609.35855(2026-09-24)· benchmark systems
  • 将被拒绝的候选配置转化为后续优化的 stepping stones,而非丢弃,避免重复陷入相同失败模式。对 prompt/harness/代码自动优化流程有启发。

8. Hebero: GPU-Parallel Heterogeneous Multi-Task RL Benchmark

  • HF Daily · arXiv 2606.03335(2026-10-05)· votes: 8 · benchmark systems
  • GPU 并行 Isaac Lab 基准,40 个异构机械臂任务联合训练评测。DGPO 方法处理稀疏奖励与有限演示。偏向机器人,agent 评测可参考其多任务评估范式。

📦 候选 JSON

/shared/research-kb/inbox/tom/_candidates/2026-10-11-agent-rag-longcontext-candidates.json

ℹ️ 备注

  • Substack:本轮搜索未发现匹配的近期高质量 Substack 研究帖(搜索词:agent RAG long context 2026),Cobus Greyling 帖子为旧文,本次未纳入。
  • CSDN:本次无适用场景(未涉及版本/命令/源码/复现经验内容),未调用。
  • 锁已持有,TTL 1500s,任务正常结束。