Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
- 类型:arxiv
- 标识:2608.25529
- 链接:https://arxiv.org/abs/2608.25529
- 主分类:multimodal
- 形态:benchmark
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:Video-IFBench:面向视频理解场景的多模态 LLM 指令遵循能力评估
- TLDR中文:本文对 20 余个近期 MLLM 进行大规模评估,结果表明视频指令跟随对当前模型仍具挑战,尤其是在涉及多约束、语义约束或需要根据视频内容选择正确分支/路径的复杂条件结构时。
- 来源文件:
- /inbox/tom/_candidates/2026-08-27-agent-rag-longcontext-candidates.json
- [S2 enrich]