论文
49 张论文卡片 · 多模态 · 评测集
人类视觉是一个闭环:注视点不断被中间假设而非单一快照持续重定向。数十年的心理物理学与认知科学研究表明,主动观察对多种任务至关重要。当代多模态大语言模型 (MLLM) 是否进行主动观察,是一个现有视觉语言基准无法回答的经验问题。我们提出 ActiveVision,一个使 MLLM 主动观察可度量的基准,包含 3 个类别共 17 个任务,任务设计强制进行重复视觉感知……Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception
提出 ENTRAP-VL(ENTRainment Assessment Probe for Vision and Language),一个由人工策展的 1,500 条数据的数据集,涵盖八个类别,按一个跨双轴的分类体系组织,并划分为文本诱发流和视觉诱发流。ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes and split into a textual-entrainment stream and a visual-entrainment stream, is introduced.
本研究考察是否存在一条实用路径,在没有坐标界面、且无需高成本区域级监督的条件下提升归因效果,并指出了这样一条可行路径。A study investigates whether there is a practical path to improve attribution without a coordinate interface and without costly region-level supervision, and indicates a practical path to improve attribution without a coordinate interface and without costly region-level supervision.
本文介绍 CLBench-V,一个多模态上下文学习 benchmark,围绕三个维度组织任务——上下文 grounding、新信息应用与新知识学习——以解决定位上下文使用失效位置的难题。This work introduces CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning.
对代表性闭源与开源多模态模型的评测表明,视觉推理强依赖于模型与环境,没有任何单一设置能在所有任务上持续占优。Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.
提出 AVE-Agent,一种模块化 agent 框架,将复杂指令分解为相互依赖的子任务,并通过自我反思与评估器反馈迭代改进编辑结果,在联合编辑中提升指令执行、保真度保持以及音视频对齐,同时保持有竞争力的感知质量。AVE-Agent is proposed, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback, and improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.
这些结果将视觉地理定位确立为场景文本仲裁的连续诊断手段,并提供了一个受控框架,用于评估 MLLMs 如何解决冲突的多模态证据。These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.
提出 FaceVid-Forensics-100K,一个大规模深伪视频数据集,包含 100,000 个视频,涵盖 33 种合成方法,覆盖换脸、表情重演与全脸合成;同时提出一个多智能体取证推理框架,由四个领域专家 Agent 分别从四个角度独立分析伪造线索。FaceVid-Forensics-100K is introduced, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, and a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives.
CLIP-CC-Bench 为长视频描述提供了一个实用的评估框架,填补了现有短片段和仅 QA 基准的空白,并通过评分者间一致性(inter-judge agreement)与 bootstrap 排序稳定性量化该协议的内可靠性。CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks and quantifying the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability.
H2R-Bench 提供了一个系统性诊断框架,用于评估视频世界模型能否跨越 human-to-robot 具身差距,并将人类操作观测转化为以机器人为中心的训练资源。H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.
本工作提出将"图像中的科学概念理解"作为长期基准目标,未来方向涵盖更广泛的领域与图表类型、上下文与跨文档综合、假设评估、出处溯源、不确定性、反事实基础,以及开放式多模态研究。This work proposes “scientific conceptual understanding from images” as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research.
认识到多样化的多参考任务共享一组共同的原子操作,本文形式化了四个算子:Anchor、Disentangle、Apply 和 Compose,并构建了 TRACE-Bench,包含约 1,600 个跨 slot 数量 1–8 的评估用例。Recognizing that diverse multi-reference tasks share a common set of atomic operations, this work formalizes four operators: Anchor, Disentangle, Disentangle, Apply, and Compose, and constructs TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8.