Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
- 类型:arxiv
- 标识:2608.25529
- 链接:https://arxiv.org/abs/2608.25529
- 主分类:multimodal
- 形态:benchmark
- TLDR:Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understandi
- 副分类:evaluation
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-08-27-agent-rag-longcontext-candidates.json