Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

  • 类型:arxiv
  • 标识:2608.15869
  • 链接:https://arxiv.org/abs/2608.15869
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos, is introduced, suggesting that explicit pixel-space generation at inference time may not be necessary for effective proactive video reasoning.
  • 副分类:llm-infra
  • 待LLM分类:否
  • 标题中文:Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
  • TLDR中文:提出 Internalized Visual Thinking (IVT),一种在无标签视频上联合优化文本预测与下一 embedding 预测的后训练框架,表明在推理时显式的像素级生成对有效的主动视频推理并非必要。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-19-agent-rag-longcontext-candidates.json
  • [S2 enrich]