研究库 论文知识库
Papers · organized/paper_cards

论文

49 张论文卡片 · 多模态 · 评测集

开放获取 全部 绿色 · 1640
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
SemComp-Bench:视频生成中的语义任务完成度基准评测
arXiv:2608.17426 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

在代表性视频生成模型上的实验表明,在保持参考图像中任务相关语义一致性的同时实现预期结果仍然具有挑战性Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.

Hydra-0: Action Flow for Generalist World Modeling and Control
Hydra-0:面向通用世界建模与控制的动作流
arXiv:2608.18077 多模态 评测集 OA · 绿色 被引 4 · S2

我们提出 Hydra-0,一种以动作流为条件的通用世界模型,将机器人动作表示为像素运动。这种共享的视觉接口使得跨具身、任务、环境和视频生成 backbone 的通用世界建模与控制成为可能,学习动作在不同场景下的后果。我们的最佳配置相比动作条件 baseline,机器人运动误差降低 90.4%,物体运动误差降低 60.2%,同时支持零样本组合与数据高效适配。在 RoboLab 基准上,Hydra-0 在 replayeWe introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replaye

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
从看见到行动:智能眼镜作为第一人称智能平台
arXiv:2608.24877 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本综述首次以统一框架系统研究智能眼镜,形式化第一人称数据流与受限任务效用,并提出覆盖采集、反应式感知、上下文辅助、持续状态、受控行动与具身耦合的 L0–L5 框架。This survey is the first to systematically study smart glasses through a unified framework, formalizing first-person data flow and constrained task utility, and introducing an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling.

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
LAION-BVD:一个用于多模态预训练的千万小时级开放视频数据集
arXiv:2608.24845 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 LAION-BVD,面向多模态学习的大规模开放视频数据集,包含从 CommonCrawl 收集的 1.3B 条平台特定视频 URL,并通过抽取场景切换帧,将视频帧作为图文数据的替代来源加以探索。This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Video-IFBench:面向视频理解场景的多模态 LLM 指令遵循能力评估
arXiv:2608.25529 多模态 评测集 被引 0 · S2

本文对 20 余个近期 MLLM 进行大规模评估,结果表明视频指令跟随对当前模型仍具挑战,尤其是在涉及多约束、语义约束或需要根据视频内容选择正确分支/路径的复杂条件结构时。This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
GUI-Primitives:诊断视觉语言 GUI grounding 中的空间推理失败
arXiv:2608.21832 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

计算机使用 Agent 将自然语言指令 grounding 到截图中以定位界面元素,但现有基准无法隔离模型是否将关系语言正确绑定到对应元素。我们提出 GUI-Primitives,一个包含 994 个条目的基准,由跨七种空间关系(左右、上下、包含、对齐、邻近、列表序数、遮挡)的对比指令对组成。每对保持截图和锚点不变,仅改变关系表达,使正确目标在两个指定候选之间切换。五位标注者……Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
PaperBanana-Interact:基于多轮人类反馈的科学图表精修
arXiv:2608.30241 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MTPaperBananaBench,一个面向多轮图表生成的基准,包含 292 张图像和 3,518 条用户需求标注,并引入 PaperBanana-Interact,一个通过内部 critique-and-refine 循环来优化图表的多智能体系统。MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
RoboSPA:VLA 模型能否超越简单场景与短时序任务?
arXiv:2609.05324 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布大规模机器人操控数据集与基准,用于诊断 VLA 模型的具身推理能力,并将 RoboSPA 确立为面向更强、更可靠、更具泛化性具身 Agent 的挑战性诊断基准。A large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models, and establishes RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents.

TempCloze: Can Video-LLMs Identify the Missing Middle?
TempCloze:Video-LLM 能否识别缺失的中间片段?
arXiv:2609.01515 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

Video-LLM 的时间推理基准常以语言为中介,留有来自选项措辞、答案相关性或语言先验带来的语言捷径空间,因此提出 TempCloze,一个用于评估 Video-LLM 视觉时间推理能力的视频完形填空基准。Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors, so TempCloze is introduced, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
arXiv:2609.10895 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

对突发物理危险的反应既是具身智能的重要检验,也是将多模态大语言模型 (MLLMs) 部署为家庭机器人决策核心的硬性要求;ReactHuman 是首个面向类人反应式决策的物理驱动基准。Reacting to sudden physical hazards is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision core of household robots, and ReactHuman is the first physics-grounded benchmark for human-like reactive decision-making.

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
MiniMax-H3 能否对物理世界进行推理?一种全模态生成模型的评估
arXiv:2609.18323 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个围绕物理世界推理四个互补维度构建的综合评估框架,并构建了一系列多样化新任务,要求模型跨模态整合互补信息。This work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning, and constructs a diverse set of novel tasks that require models to integrate complementary information across modalities.

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
OmniVChat:面向原生音视频对话的合成、基准与训练
arXiv:2609.21465 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

OmniVChat-Studio,一种用于合成单轮与多轮音视频对话的多 Agent 数据引擎;以及 OmniVChat-RL,一种在 OmniVChat 中同时针对回复正确性、效率与风格设计奖励的强化学习奖励方案。OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues and OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat.

Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
Tri-PvP:通过感知-命题证据冲突揭示 Omni-Modal LLM 的模态偏差
arXiv:2609.06011 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Tri-PvP,一个包含 8,000 样本、跨越视觉、音频与文本的三模态冲突基准,揭示了证据形式偏倚的系统性不对称:模型在视觉上更偏向感知信号,而在音频上更偏向命题性信号。Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, is introduced, revealing a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio.

OmniEcho: Spatial Audio Understanding for Embodied Agents
OmniEcho:面向具身智能体的空间音频理解
arXiv:2609.23407 多模态 评测集 OA · 绿色 被引 1 · S2

本文提出 OmniEchoBench,一个面向空间音视频感知与音-视-语言导航的统一基准,以及一个空间感知的全模态模型,该模型在预训练语义音频通路之外引入 FOA 空间编码器。OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation, and a spatially aware omni-modal model, which introduces an FOA spatial encoder alongside a pretrained semantic audio pathway.

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
难视觉,易视觉:GPT-6 Astra 揭示的计算机视觉全貌
arXiv:2609.35718 多模态 评测集 OA · 绿色 被引 1 · S2

本文勾勒了一幅计算机视觉版图:在其中,越来越复杂的视觉任务可通过通用接口访问,而精确且对保真度敏感的感知仍是重要前沿。A changing landscape of computer vision is mapped in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
[标题中文] APM-Bench:面向自我中心流式视频助手的跨会话持久记忆基准
arXiv:2609.37559 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

APM-Bench 将真实场景的流式交互重构为多会话生命轨迹,并揭示了一个清晰的效用—延迟—存储权衡:现有方法仍难以同时实现可靠的长程记忆、低开销以及跨会话有效的主动协助。APM-Bench is introduced, which reformulates real-world streaming interaction as multi-session life trajectories and reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions.

Multimodal 补充候选
arXiv:2508.17398 多模态 评测集 被引 4 · S2

本文提出 DashboardQA,这是首个明确设计用于评估视觉-语言 GUI Agent 对真实世界仪表板理解与交互能力的基准,结果表明交互式仪表板推理对所有受评估的 VLM 而言都是一项具有挑战性的任务。DashboardQA is introduced, the first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards, and indicates that interactive dashboard reasoning is a challenging task overall for all the VLMs evaluated.

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Ego2Act:评估自我中心视频生成中的目标导向操作
arXiv:2610.01092 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Ego2Act:一个面向目标的基准,包含来自 110 个真实日常任务的 2.640 段视频,覆盖不同的物体杂乱度与多步复杂度;同时提出 Ego2ActJudge,一种无参考评估流水线,在任务完成度与物理合理性评估上与人类共识的对齐效果优于相关基线。Ego2Act is introduced, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity, and Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines.

World Embedding Benchmark
World Embedding Benchmark(世界嵌入基准)
arXiv:2610.03632 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 World Embedding Benchmark,包含来自 80 个族的 8,000 个受控仿真案例,涵盖流体力学、固体力学、动力学以及光学与电磁学,旨在强调需要联合评估物理一致性与物理属性可恢复性,并展示物理表示对改进视频生成的效用。The World Embedding Benchmark is introduced, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics&electromagnetism, to highlight the need to evaluate physical alignment and property recoverability jointly and demonstrate the utility of physical representations for improving video generation.

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
HyperBrowseComp:面向网页浏览 Agent 的多语言多模态压力测试
arXiv:2610.03574 多模态 评测集 被引 0 · S2

提出了 HyperBrowseComp,一个多语言、多模态的浏览基准,包含由母语或高水平使用者编写并经人工验证的、跨 13 种语言的 423 个问题,难度源于在开放网络上发现并关联证据。HyperBrowseComp is introduced, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers, with difficulty arising from discovering and connecting evidence on the open web.

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models
双臂视觉-语言-动作模型中的按臂组合泛化
arXiv:2610.06184 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ACG-Bench,一个面向逐臂组合泛化(arm-wise Compositional Generalization)的基准,为超越固定训练流程的双臂策略设计提供实证指导,考察了 arm-token grouping、技能专属 LoRA 适配器(SkillLoRA)以及 arm-wise attention(AWA),凸显了技能条件化参数与注意力结构的互补性。This work introduces ACG-Bench, a benchmark for arm-wise Compositional Generalization that provides empirical guidance for designing dual-arm policies that generalize beyond fixed training routines, and examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure.

OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
OmniCapBench:面向细粒度音视频描述的深度结构化评估框架
arXiv:2610.12458 多模态 评测集

多模态大语言模型(MLLMs)正快速发展以实现持续的音视频推理,这迫切需要能够揭示其能力上限的评估。音视频描述是一项理想的诊断任务,但现有 benchmark 面临耦合的权衡:整句评分覆盖全面但缺乏定位,局部探针定位精确但缺乏覆盖,且无约束的 LLM 评判器带来不稳定性。我们提出 OmniCapBench(Omni-Video Caption Benchmark),将音视频描述评估重构为深度结构化的诊断框架Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic fr

DataComp-VLM: Improved Open Datasets for Vision-Language Models
DataComp-VLM:面向视觉-语言模型的改进开源数据集
arXiv:2606.28551 多模态 评测集 MPG.PuRe (Max Planck Society) OA · 绿色 被引 1 · S2

数据混合(而非过滤)是构建高质量训练数据集的关键:以指令型数据为主的混合在扩展时优于以描述型数据为主的混合,且规模越大优势越明显。It is found that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales.

Discrete Diffusion Language Models for Interactive Radiology Report Drafting
用于交互式放射学报告起草的离散扩散语言模型
arXiv:2607.01436 多模态 评测集 OA · 绿色 被引 1 · S2

本文适配了一款专家混合扩散语言模型 DiffusionGemma-26B,并在医学视觉问答数据集上,使用相同的 LoRA 配置将其与同规模的自回归模型 Gemma-4-26B 进行基准对比,由对冗长度鲁棒的 LLM 裁判打分。This work adapts a mixture-of-experts diffusion language model, DiffusionGemma-26B, and benchmark it against its same-size AR sibling Gemma-4-26B under an identical LoRA recipe on medical visual question answering datasets, scored by a verbosity-robust LLM judge.

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
打破失败级联:面向医学多模态推理的步骤感知强化学习
arXiv:2606.31825 多模态 评测集 OA · 绿色 被引 2 · S2

在四个多模态 LLM backbone 上,MRPO 始终优于标准 GRPO 与一项最新 RL baseline,并在 Qwen3-VL-8B-Thinking 上以 4.59 分超越规模远大于它的医学 MLLM(如 HuatuoGPT-Vision-34B)。Across four multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3-VL-8B-Thinking even surpasses substantially larger medical MLLMs such as HuatuoGPT-Vision-34B by 4.59 points.

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
ABACUS:适配统一基础模型以桥接图像计数理解与生成
arXiv:2606.23835 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

ABACUS 提出三项贡献:与基于多头自注意力分解得到的目标性图相结合的密度感知自适应缩放,用于在空间上锚定计数预测;通过 GRPO 训练的边界感知计数策略,配合嵌套的局部、边界与全局奖励,以消除裁剪边界处的过度与不足计数。ABACUS introduces three contributions: density-aware adaptive zooming paired with an objectness map from multi-head self-attention decomposition to spatially ground count predictions; a boundary-aware count policy trained via GRPO with nested local, boundary, and global rewards to eliminate over- and undercounting at crop boundaries.

Video-Oasis: Rethinking Evaluation of Video Understanding
Video-Oasis:重新审视视频理解评估
arXiv:2603.29616 多模态 评测集 OA · 绿色 被引 2 · S2

该审计揭示现有基准样本中 55% 可在无视觉输入或时序上下文的情况下被解决,并提出 Video-Oasis,一个用于系统性审计现有视频理解基准的可持续诊断套件。This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context, and introduces Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks.

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
KeyFrame-Compass:迈向关键帧条件视频生成的综合评估
arXiv:2607.14202 多模态 评测集 OA · 绿色 被引 1 · S2

本文将关键帧执行拆解为存在性、保真度、时序顺序、定位、持续性与唯一性六个互补指标,并通过结合专用感知模型的、基于证据的 MLLM 判断来评估整体视频质量。This work decomposes keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
MultiRef-Compass:迈向多参考音视频生成任务的综合评估
arXiv:2607.14189 多模态 评测集 OA · 绿色 被引 6 · S2

本文提出 MultiRef-Compass,一个面向 MR2AV 生成的统一基准,将自动指标与引入复判增强的 MLLM-as-a-Judge 框架相结合,实现对感知保真度与参考条件合成能力的可扩展、可审计评估。MultiRef-Compass is introduced, a unified benchmark for MR2AV generation that integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition.

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
VIABench:面向视障辅助任务的盲人视频综合基准
arXiv:2607.14660 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 VIABench,一个专为评估 MLLM 在视障辅助(VIA)场景中表现而设计的综合视频基准,采用视障人士(VIIs)自行录制或分享的第一人称视频,并提出一套严格的评测流水线,同时支持在线(实时)与离线设置。VIABench is introduced, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves, and proposes a rigorous benchmarking pipeline that supports both online (real-time) and offline settings.

ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
ScanNet:富含标注的室内场景三维重建
arXiv:1702.04405 多模态 评测集 OA · 绿色 被引 6002 · S2

本文推出 ScanNet,一个 RGB-D 视频数据集,包含 1513 个场景中的 250 万视角,标注有三维相机位姿、表面重建与语义分割,并表明使用该数据可在多项三维场景理解任务上取得 SOTA 性能。This work introduces ScanNet, an RGB-D video dataset containing 2.5M views in 1513 scenes annotated with 3D camera poses, surface reconstructions, and semantic segmentations, and shows that using this data helps achieve state-of-the-art performance on several 3D scene understanding tasks.

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
ChatGPT 在推理、幻觉与交互性方面的多任务、多语言、多模态评估
arXiv:2302.04023 多模态 评测集 OA · 绿色 被引 1839 · S2

研究发现 ChatGPT 在大多数任务上以零样本学习优于其他 LLM,在部分任务上甚至超过微调模型,并且对非拉丁文字语言的理解能力优于生成能力。It is found that ChatGPT outperforms LLMs with zero-shot learning on most tasks and even outperforms fine-tuned models on some tasks and is better at understanding non-Latin script languages than generating them.

Matterport3D: Learning from RGB-D Data in Indoor Environments
Matterport3D:基于室内 RGB-D 数据的学习
arXiv:1709.06158 多模态 评测集 OA · 绿色 被引 2726 · S2

本文介绍 Matterport3D,一个大规模 RGB-D 数据集,包含来自 90 个建筑物级场景共 194,400 张 RGB-D 图像的 10,800 个全景视图,可支持多种监督与自监督计算机视觉任务,包括关键点匹配、视角重叠预测、由彩色图像预测法线、语义分割和区域分类。Matterport3D is introduced, a large-scale RGB-D dataset containing 10,800 panoramic views from 194,400RGB-D images of 90 building-scale scenes that enable a variety of supervised and self-supervised computer vision tasks, including keypoint matching, view overlap prediction, normal prediction from color, semantic segmentation, and region classification.

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
SVR-R1:在强化学习中通过自验证自举多模态推理
arXiv:2607.10966 多模态 评测集 OA · 绿色 被引 1 · S2

本文提出自验证推理器 SVR-R1,一种多轮强化学习框架,将模型自身的验证转化为多模态推理的学习信号,提供了一种简洁而有效的多模态推理自举方案。Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning, is introduced, offering a simple yet effective recipe for bootstrapping multimodal reasoning.

Can Multimodal Large Language Models Understand OCT?
多模态大语言模型能理解OCT吗?
arXiv:2607.16609 多模态 评测集 OA · 绿色 被引 3 · S2

OCT-Bench能够对MLLM进行全面且细粒度的评估,为识别能力瓶颈和推进临床可信的OCT理解奠定基础。OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.

GigaChat Audio: Time-aware Large Audio Language Model
GigaChat Audio:时间感知的大音频语言模型
arXiv:2607.10387 多模态 评测集 OA · 绿色 被引 1 · S2

本文提出一个时间感知的音频 LLM,能够基于大规模合成监督(来自级联 pipeline)在长达 120 分钟的输入上回答带有显式时间戳的问题,并在短时长和长时长 benchmark 上取得强劲的时间定位准确率。This work presents a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input using large-scale synthetic supervision from a cascaded pipeline and achieves strong temporal-grounding accuracy on short and long benchmarks.