2026-09-30 周三多模态文献总结(flyP 周报)

今日主题

多模态:图像生成、视频生成、音频生成、视觉语言模型(VLM)、多模态推理与评估;本周重点关注“联合音视频 RL / 视觉感知瓶颈 / VLM 长文档 / 量化幻觉 / 世界模型内部表征”。

检索来源

  • arXiv cs.CV / cs.CL / cs.SD / cs.MM / cs.GR / cs.AI / cs.LG new + pastweek(2026-09-23 → 2026-09-30)
  • Hugging Face 论文页与 Papers-with-Code 提及
  • ICML 2026 / ICLR 2026 / NeurIPS 2026 / CVPR 2026 / EACL 2026 / EMNLP 2026 已接收论文集
  • Substack 高价值 newsletter:The Living Edge(Philip Bankier,#40 + #41 双周回顾)、Jakob Nielsen PhD 2026 预测、Future AGI
  • awesome-video-diffusions、Video-Generation-arxiv-daily、PRIV-Creation/Awesome-Controllable-T2I-Diffusion-Models
  • CSDN:本轮未发现符合“版本+命令+复现/排障”高价值工程稿,不收录

新增候选概览(2026-09-23 → 2026-09-30 arXiv 与公开资料)

ID 标题 类别 提交/接收 备注
2609.29816 AV-GRPO: Modality-Anchored Decoupling Diffusion RL for Joint Audio-Video Generation audio-video / RL 2026-09-29 (new) 把 GRPO 引入扩散联合音视频;22 页大稿
2609.30210 The Alignment Illusion in Multimodal Large Language Models VLM / evaluation NeurIPS 2026 揭示 MLLM 在 benchmark 上的“伪对齐”现象
2609.29607 STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video LLMs VLM / benchmark 2026-09-29 Video-LLM 的对象级时空监控
2609.29999 GHOST-Q: Studying Grounding Hallucinations Under Same-score Trade-offs in Quantized VLMs VLM / quantization ICASSP 2027 投稿 量化 VLM 的 grounding hallucination
2609.29933 An Empirical Study of VLM Pipelines for Long-Document QA VLM / long-doc EMNLP 2026 Industry 22 页 VLM 长文档 QA 流水线对比
2609.34826 WM-VLM: Probing Internal World Models for Interleaved Text-Visual Reasoning VLM / world model 2026-09 验证 VLM 能否用生成视觉状态辅助推理
2609.31652 Open-Qwen-Music: An Auditable Framework for LLM-Based Music Composition and Diffusion Rendering audio / music 2026-09 LLM 作曲 + diffusion 渲染,可审计
2609.34901 Domain-Incremental Learning for Generative Speech Enhancement audio / speech 2026-09 域增量生成式语音增强
2609.30912 Music Source Separation via Stem Discovery audio / separation ICASSP 2027 stem 发现式音源分离
2609.30187 Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures video / 4D 2026-09-29 数据集:自我+他人视角 4D 人体网格
2609.29508 Evaluation of Multi-Turn Consistency in LLM Agents (failure-rationale taxonomy) agent / eval ICLR 2026 Workshop 与多模态 agent 间接相关
VLM-CapCurriculum (UCSC-VLAA, ICML 2026) From Seeing to Thinking: Decoupling Perception and Reasoning Improves VLM Post-Training VLM / curriculum ICML 2026 “看错再长 CoT 也救不回来”,staged curriculum
LaViDa (NeurIPS 2025/2026 讨论) A Large Diffusion Language Model for Multimodal Understanding VLM / discrete diffusion NeurIPS 2025 poster 离散扩散 VLM,可控并行解码
Paper-Notes-en / SafeGRPO (CVPR 2026) Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization VLM / safety CVPR 2026 Qwen3-VL-4B/8B-Thinking + GRPO + verl
ICML 2026 PAPO Perception-Aware Policy Optimization for Multimodal Reasoning VLM / RL ICLR 2026 改 KL 目标而非 rollout/reward,4.4-17.5% 提升
Substack · Last Week in Multimodal AI #40 / #41 周更 newsletter survey 2026-09 中下旬 #40 关键词:Qwen3-VL-Embedding、e5-omni、HY-Video-PRFL;#41:Vision Model Reality、Agent Memory
Substack · Future AGI “Multimodal AI in 2026” 行业综述 survey 2026 多模态“production-ready”叙事,含跨模态注意力说明
Jakob Nielsen PhD · 2026 Predictions 行业预测 survey 2026 “text-LLM 时代结束,Large World Models 到来”;提示单模态公司可能被收购
DigitalApplied · Multimodal AI Benchmarks 2026 基准评论 evaluation 2026-04 MMMU-Pro 已饱和(4 模型 81-83%),需按模态挑模型

说明:所有条目仅做中文摘记 + 可信度判断 + 链接引用;不复制原文。

必读 3-5 篇

  1. AV-GRPO(arXiv 2609.29816):把 GRPO 引入“联合音视频扩散”的 modality-anchored decoupling。22 页大稿,做音频/视频联合生成的 RL 微调,是对“Thinking with Video”/“audio-visual reasoning”路线的工程化深化。建议精读 + 跟踪是否开源。
  2. The Alignment Illusion in Multimodal LLMs(arXiv 2609.30210,NeurIPS 2026):提出 MLLM 在多个 alignment benchmark 上的“伪对齐”假象——指标好看但实际行为未对齐。属于评估/对齐方向的反方审稿线索,必读。
  3. VLM-CapCurriculum(UCSC-VLAA,ICML 2026,project page):把 VLM 后训练拆成“视觉感知 → 文本推理 → 视觉推理”三阶段,能在 4 个 backbone 上让推理链缩短 20.8% 同时涨点。“longer thinking ≠ better”的实证,与 Alignment Illusion 形成呼应。
  4. STRAND(arXiv 2609.29607):Video LLM 的对象级时空监控 benchmark + 方法。视频 agent / 长视频理解方向必读,作为“object permanence 类”评测的补集。
  5. Open-Qwen-Music(arXiv 2609.31652):LLM 作曲 + diffusion 渲染的“可审计”框架。可与 UAT、DreamAudio 并列作为“生成式音频的可解释性”主线;评估音乐生成的 auditability 是稀缺视角。

高价值技术文章

  • The Living Edge · Last Week in Multimodal AI #40 / #41(Substack,Philip Bankier):#40 关键词 = Qwen3-VL-Embedding + e5-omni 跨模态检索、HY-Video-PRFL(视频扩散自反馈 + 56% 动效提升)、PointWorld-1B、RoboVIP、Web World Models、cbottle、LTX-2;#41 主题 = “Vision Model Reality + Agent Memory Upgrade”。建议作为本周主 newsletter 线索源长期订阅。
  • Jakob Nielsen PhD · 2026 Predictions comic(Substack):把多模态/世界模型列为 Prediction 11(“text-LLM 时代结束,Large World Models 到来”),Prediction 12 预测单模态公司(Flux / Ideogram / Midjourney / Leonardo / Reve)将被多模态实验室收购。行业风向标,仅作背景。
  • Future AGI · Multimodal AI in 2026(Substack):以 GPT-5 为锚点讲多模态进入 production-ready;含 cross-attention 解释、医学影像+临床文本案例。适合做新人科普,不建议作为唯一信源。
  • DigitalApplied · Multimodal AI Benchmarks 2026:MMMU-Pro 已饱和(GPT-5.5 / Gemini 3 / Claude Opus 4.7 / Qwen 3.5 Omni 全在 81-83%);提出“按模态挑模型”而非看头条分数。与 Alignment Illusion 形成互相印证。
  • VLM-CapCurriculum 项目页(UCSC-VLAA):图示“长 CoT 救不回错看”与“staged curriculum”,可视化做得好;可截图作为内部周会材料。
  • Paper-Notes-en · SafeGRPO 笔记(CVPR 2026):给出 Qwen3-VL-4B/8B-Thinking + verl + GRPO 的训练细节与基准分数(SIUO / MOSSBench / refusal rate),可作为 GRPO 多模态化的工程模板参考。
  • CSDN 本轮未发现符合“版本+命令+复现/排障”标准的高价值工程稿,不收录。

分类标签

multimodal / video generation / image generation / audio generation / VLM / multimodal reasoning / audio-video / RL / GRPO / perception / evaluation / benchmark / quantization / world model / safety / long-document / music / survey / negative-result / reproduction

是否建议精读 / 反方审稿 / 主题页更新

  • 建议精读:AV-GRPO、Alignment Illusion、VLM-CapCurriculum、STRAND、WM-VLM。
  • 建议作为反方审稿(看质疑点):
  • AV-GRPO:22 页大稿,是否在同步语音/音乐上做了 alignment?目前只给“joint audio-video”客观分数,缺乏人评。
  • Alignment Illusion:是否覆盖闭源模型(GPT-5 系、Gemini 3、Claude Opus 4.7、Qwen 3.5 Omni)的“伪对齐”案例?需要确认是否仅在开源模型上验证。
  • VLM-CapCurriculum:“staged curriculum”依赖感知标注成本,需要量化降本比;论文目前只强调效果。
  • GHOST-Q:在同分权衡下测 grounding hallucination,metric 设计是否会被模型针对性刷分?
  • STRAND:对象级时空监控的标注一致性需要披露 inter-annotator agreement。
  • 建议主题页更新:
  • “VLM 评估 / 对齐” 加入 Alignment Illusion、PAPO(ICLR 2026)、SafeGRPO(CVPR 2026)、MMMU-Pro 饱和警示(DigitalApplied);
  • “视频 / 时空理解” 加入 STRAND、WM-VLM、Ego-Exo4D Meshes;
  • “音频生成” 加入 Open-Qwen-Music、Stem Discovery 音源分离、UAT、DreamAudio;
  • “联合音视频 RL” 单列新主题,加入 AV-GRPO;
  • “多模态世界模型” 加入 WM-VLM、PointWorld-1B(来自 #40)。

建议写入文件路径

  • 本周报草稿:/shared/research-kb/inbox/flyp/2026-09-30-multimodal-weekly-digest.md(即本文件)
  • 同步根路径:research-kb/digests/2026-09-30_multimodal.md(由同步任务统一合并)
  • 候选登记:research-kb/registry/papers.jsonl 追加下列行(不入 commit,由合并任务串行追加):
{"id":"2609.29816","title":"AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation","venue":"arXiv","date":"2026-09-29","tags":["audio-video","diffusion","RL","GRPO"],"note":"22 pages; joint AV RL post-training"}
{"id":"2609.30210","title":"The Alignment Illusion in Multimodal Large Language Models","venue":"arXiv (NeurIPS 2026)","date":"2026-09-29","tags":["VLM","evaluation","alignment","negative-result"],"note":"benchmark-good but behavior-not-aligned"}
{"id":"2609.29607","title":"STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models","venue":"arXiv","date":"2026-09-29","tags":["VLM","video","benchmark","spatio-temporal"],"note":"object-centric video-LLM eval"}
{"id":"2609.29999","title":"GHOST-Q: Studying Grounding Hallucinations Under Same-score Trade-offs in Quantized VLMs","venue":"arXiv (ICASSP 2027)","date":"2026-09-29","tags":["VLM","quantization","hallucination","evaluation"],"note":"trade-off aware hallucination under quantization"}
{"id":"2609.29933","title":"An Empirical Study of VLM Pipelines for Long-Document QA","venue":"arXiv (EMNLP 2026 Industry)","date":"2026-09-29","tags":["VLM","long-document","QA","systems"],"note":"22-page pipeline survey"}
{"id":"2609.34826","title":"WM-VLM: Probing Internal World Models for Interleaved Text-Visual Reasoning","venue":"arXiv","date":"2026-09","tags":["VLM","world model","interleaved reasoning"],"note":"generated visual states aid reasoning"}
{"id":"2609.31652","title":"Open-Qwen-Music: An Auditable Framework for LLM-Based Music Composition and Diffusion Rendering","venue":"arXiv","date":"2026-09","tags":["audio","music","diffusion","LLM"],"note":"LLM + diffusion render; auditable"}
{"id":"2609.34901","title":"Domain-Incremental Learning for Generative Speech Enhancement","venue":"arXiv (eess.AS)","date":"2026-09","tags":["audio","speech","continual-learning"],"note":"domain-incremental speech enhancement"}
{"id":"2609.30912","title":"Music Source Separation via Stem Discovery","venue":"arXiv (ICASSP 2027)","date":"2026-09","tags":["audio","music","separation"],"note":"stem-discovery based source separation"}
{"id":"2609.30187","title":"Ego-Exo4D Human Meshes Dataset","venue":"arXiv","date":"2026-09-29","tags":["video","4D","dataset","human-motion"],"note":"ego+exo 4D meshes"}
{"id":"VLM-CapCurriculum","title":"From Seeing to Thinking: Decoupling Perception and Reasoning Improves VLM Post-Training","venue":"ICML 2026","date":"2026","tags":["VLM","curriculum","post-training","perception"],"note":"20.8% shorter CoT, +accuracy across 4 backbones"}
{"id":"PAPO","title":"Perception-Aware Policy Optimization for Multimodal Reasoning","venue":"ICLR 2026","date":"2026-04","tags":["VLM","RL","GRPO","DAPO"],"note":"+4.4-17.5% via new KL objective; -30.5% perception error"}
{"id":"SafeGRPO","title":"Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization","venue":"CVPR 2026","date":"2026","tags":["VLM","safety","GRPO"],"note":"Qwen3-VL-4B/8B-Thinking + verl"}
{"id":"LaViDa","title":"A Large Diffusion Language Model for Multimodal Understanding","venue":"NeurIPS 2025 poster","date":"2025","tags":["VLM","discrete-diffusion","multimodal"],"note":"parallel decoding + controllable generation"}
{"id":"thelivingedge-40","title":"Last Week in Multimodal AI #40: Search Across Everything","venue":"Substack (Philip Bankier)","date":"2026-09","tags":["survey","multimodal","newsletter"],"note":"Qwen3-VL-Embedding, e5-omni, HY-Video-PRFL, cbottle"}
{"id":"thelivingedge-41","title":"Last Week in Multimodal AI #41: Vision Model Reality, Agent Memory Upgrade","venue":"Substack (Philip Bankier)","date":"2026-09","tags":["survey","multimodal","newsletter"],"note":"vision model reality check + agent memory"}
{"id":"futureagi-2026","title":"Multimodal AI in 2026: What's Happening Now and What's Coming Next","venue":"Substack (Future AGI)","date":"2026","tags":["survey","industry"],"note":"GPT-5 anchored cross-attention explainer"}
{"id":"digitalapplied-bench-2026","title":"Multimodal AI Benchmarks 2026: Vision, Audio, Code","venue":"DigitalApplied blog","date":"2026-04","tags":["benchmark","evaluation","survey"],"note":"MMMU-Pro saturated 81-83% across 4 frontier models"}

待人工确认的问题

  1. AV-GRPO(2609.29816)是否提供 checkpoint / 代码?摘要中未注明,建议打开论文 + project page 后再标 reproduction。
  2. Alignment Illusion(2609.30210)实验对象是否覆盖闭源模型(GPT-5.5、Gemini 3、Claude Opus 4.7、Qwen 3.5 Omni)?目前 DigitalApplied 与论文均强调 MMMU-Pro 饱和,但需要看是否同时给闭源对位结果。
  3. VLM-CapCurriculum(ICML 2026)的“staged curriculum”成本曲线(标注成本 vs 准确率增益)是否披露?若仅有最终分数,需注明 “no-cost-analysis” 标签。
  4. SafeGRPO(CVPR 2026)使用了 Qwen3-VL-Thinking + verl,是否已在 Qwen3.5 / InternVL / InternLM-XComposer 上做迁移?是否仅在 Qwen 系列验证,需在审稿意见里标出。
  5. The Living Edge #41(2026-09 周中)的具体发文日期未在搜索结果中直接显示,建议在订阅页面核对发表时间后再决定纳入哪一周的周报。
  6. 是否需要把 “perception-aware post-training”(VLM-CapCurriculum + PAPO + SafeGRPO)合并为新的主题页 “VLM Post-Training 2.0”,以避免分散到 curriculum / RL / safety 三个不同主题?
  7. Substack 来源本周以 The Living Edge 为主(#40 + #41),Jakob Nielsen PhD 与 Future AGI 是否要纳入“季度趋势回顾”而非每周多模态周报?