Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

  • 类型:arxiv
  • 标识:2608.03979
  • 链接:https://arxiv.org/abs/2608.03979
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:Video-DR is introduced, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval, enabling autonomous exploration that breaks the imitation-learning ceiling.
  • 副分类:agent
  • 待LLM分类:否
  • 标题中文:Video-DeepResearch:迈向下一代多模态深度研究 Agent
  • TLDR中文:提出 Video-DR,采用解耦的感知-探索流水线与分阶段工具解锁,强制在 web 检索前进行充分的跨帧视觉定位,实现突破模仿学习上限的自主探索。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-05-agent-rag-longcontext-candidates.json
  • [S2 enrich]