Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

  • 类型:arxiv
  • 标识:2608.25529
  • 链接:https://arxiv.org/abs/2608.25529
  • 主分类:multimodal
  • 形态:benchmark
  • TLDR:Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understandi
  • 副分类:evaluation
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-08-27-agent-rag-longcontext-candidates.json