FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
- 类型:arxiv
- 标识:2608.21839
- 链接:https://arxiv.org/abs/2608.21839
- 主分类:multimodal
- 形态:method
- TLDR:Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal v
- 副分类:evaluation
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-08-27-agent-rag-longcontext-candidates.json