日期:2026-10-02 · 轮次 R2 arxiv 链接:https://arxiv.org/abs/2609.37863 video-kb 脚本原路径:/shared/video-kb/scripts/tom/2026-10-02_R2_vise-mll-strike-show-config-not-model_short-video-script.md
1. 标题与观众
- 标题:VLM 裁判 别只看分
- 目标观众:AI 评测工程师(主)· VLM / 多模态系统架构师(次)· 用 VLM 当人类标注替代品的技术产品经理(第三)
- 一句话核心观点:评测时给 VLM 裁判塞张图就漂分,但漂移不指图像语义——测的是配置本身。
- 备选标题 1:200 句三档图暴露评测失真
- 备选标题 2:评测协议不是模型是配置
3. 45-90 秒口播稿
VLM 当人类标注员就一定更强?MIST 给出答案:200 句 × 3 档图 = 600 评估配置、13 个 VLM judge + alt-test sanity check,7 strong + 6 not-passed 两组。
数字 verbatim:aligned 20.5%、misleading 19.4%、删指令基线 11.6%——两图漂移都高于基线,方向一致率 37%。
结论 verbatim:what moves a judge is that an image is there, not which of the two it is;a substitutability verdict describes a configuration as much as a model。
工程可执行:image=None 与 image=present 双列必报;ignore-the-image 指令 prompt,靠不住。
4. 镜头分镜
- 0-6s · 标题卡 屏幕文字「VLM 裁判 别只看分」,画面 = 黑底白字标题卡 + 左下角 VLM logo 简笔 + 右下角「评测协议失真 = 配置失真」限定语字幕,运动 = 标题卡轻微放大 + 限定语字幕淡入。
- 6-14s · 流程图 屏幕文字「multimodal judge 默认能看图 · 假设一定比纯文本 judge 强」,画面 = 流程图(multimodal judge 框 → 假设一定更强 框 → 撞墙图标),运动 = 流程图自左向右逐框淡入 + 撞墙图标闪烁 3 次。
- 14-24s · 信息图 屏幕文字「200 句 literal/figurative 双解读 × 3 档图 = 600 配置 · 13 VLM judges + alt-test」,画面 = 200 句表格(literal/figurative 双解读短语示例 = 2 行)+ 3 档图对照(aligned / misleading / no-image 三栏)+ 13 judge 散点图占位,运动 = 表格左滑入 → 3 档图右滑入 → 散点图渐显。
- 24-32s · 证据卡 屏幕文字「aligned 20.5% · misleading 19.4% · 删除指令基线 11.6%」,画面 = 柱状图(aligned 高 / misleading 中 / 基线低)+ abstract verbatim 字面字幕,运动 = 柱状图自下而上生长 + abstract verbatim 字面字幕下方对齐。
- 32-39s · 证据卡 屏幕文字「方向一致率 37% · 漂移 ≠ 图像语义方向」,画面 = 饼图(37% 命中 vs 63% 无方向漂移)+ 箭头指向「被忽略的图像语义方向」,运动 = 饼图顺时针分块动画 + 箭头闪烁 2 次。
- 39-46s · 标题卡 屏幕文字「what moves a judge is that an image is there, not which of the two it is」,画面 = abstract verbatim 措辞标题卡 + 限定语字幕「config = 配置 · 测的是配置不是模型」,运动 = abstract verbatim 字面字幕居中显示 + 限定语字幕下方对齐淡入。
- 46-51s · 风险卡 屏幕文字「评测时必报 image=None 与 image=present 双列 · ignore 指令无效 · logit/attn 屏蔽修复」,画面 = 红色边框风险卡 + 三条要点卡,运动 = 风险卡自上而下逐条淡入 + 边框轻微脉动。
- 51-60s · 行动清单 + closing 屏幕文字「评测时双列必报 · prompt 假装忽略忽略指令 · 中性无关图必加 · 200 句之外自行复现」,画面 = 行动清单(4 条 bullet · 黑底白字)+ closing 字幕「把判断变成行动」,运动 = 清单逐条滑入 + closing 字幕淡入收尾。