Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
- 类型:arxiv
- 标识:2608.03979
- 链接:https://arxiv.org/abs/2608.03979
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:Video-DR is introduced, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval, enabling autonomous exploration that breaks the imitation-learning ceiling.
- 副分类:agent
- 待LLM分类:否
- 标题中文:Video-DeepResearch:迈向下一代多模态深度研究 Agent
- TLDR中文:提出 Video-DR,采用解耦的感知-探索流水线与分阶段工具解锁,强制在 web 检索前进行充分的跨帧视觉定位,实现突破模仿学习上限的自主探索。
- 来源文件:
- /inbox/tom/_candidates/2026-08-05-agent-rag-longcontext-candidates.json
- [S2 enrich]