TempCloze: Can Video-LLMs Identify the Missing Middle?
- 类型:arxiv
- 标识:2609.01515
- 链接:https://arxiv.org/abs/2609.01515
- 主分类:multimodal
- 形态:benchmark
- TLDR:Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Seman
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:TempCloze:Video-LLM 能否识别缺失的中间片段?
- TLDR中文:Video-LLM 的时序推理基准通常以语言为中介,选项措辞、答案相关性或语言先验都可能导致语言捷径。为减少此类捷径,我们推出 TempCloze,一个用于评估 Video-LLM 视觉时序推理能力的视频完形填空基准。给定视频的开头和结尾片段,模型需从四个候选中识别出真正的缺失中间片段。TempCloze 包含来自七个来源的 1,521 个经过精心筛选的视频,主要为长镜头和第一人称视角视频。我们沿三个维度构建同源干扰项:语义
- 来源文件:
- /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json