WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning

  • 类型:arxiv
  • 标识:2609.34826
  • 链接:https://arxiv.org/abs/2609.34826
  • 主分类:multimodal
  • 形态:method
  • TLDR:Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasonin
  • 副分类:engineering
  • 待LLM分类:否
  • 标题中文:WM-VLM:探查用于交错式视觉-文本推理的内部世界模型
  • TLDR中文:人类常通过心智模拟视觉变换来解决空间问题,而传统视觉语言模型(VLM)主要依赖语言进行推理。我们探究 VLM 是否能通过同时基于文本和生成的视觉状态进行推理来解决空间问题。为此,我们提出 WM-VLM,为预训练 VLM 配备一个轻量级的世界模型分支,用于生成中间视觉状态。训练分两阶段:首先让模型学会生成下一视觉状态,再让其学会利用该状态进行推理。我们以程序化方式构建了空间推理……
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-06-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-10-06-agent-rag-longcontext-candidates.json