Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
- 类型:arxiv
- 标识:2608.15869
- 链接:https://arxiv.org/abs/2608.15869
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos, is introduced, suggesting that explicit pixel-space generation at inference time may not be necessary for effective proactive video reasoning.
- 副分类:llm-infra
- 待LLM分类:否
- 标题中文:Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
- TLDR中文:提出 Internalized Visual Thinking (IVT),一种在无标签视频上联合优化文本预测与下一 embedding 预测的后训练框架,表明在推理时显式的像素级生成对有效的主动视频推理并非必要。
- 来源文件:
- /inbox/tom/_candidates/2026-08-19-agent-rag-longcontext-candidates.json
- [S2 enrich]