TempCloze: Can Video-LLMs Identify the Missing Middle?

  • 类型:arxiv
  • 标识:2609.01515
  • 链接:https://arxiv.org/abs/2609.01515
  • 主分类:multimodal
  • 形态:benchmark
  • TLDR:Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Seman
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:TempCloze:Video-LLM 能否识别缺失的中间片段?
  • TLDR中文:Video-LLM 的时序推理基准通常以语言为中介,选项措辞、答案相关性或语言先验都可能导致语言捷径。为减少此类捷径,我们推出 TempCloze,一个用于评估 Video-LLM 视觉时序推理能力的视频完形填空基准。给定视频的开头和结尾片段,模型需从四个候选中识别出真正的缺失中间片段。TempCloze 包含来自七个来源的 1,521 个经过精心筛选的视频,主要为长镜头和第一人称视角视频。我们沿三个维度构建同源干扰项:语义
  • 来源文件
  • /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json