Papers · organized/paper_cards

论文

258 张论文卡片 · 多模态

开放获取 全部 绿色 · 677
From Foundation to Application: Improving VLA Models in Practice
从基础到应用:实践中改进 VLA 模型
arXiv:2607.06403 多模态 应用落地 OA · 绿色 被引 11 · S2

得益于涵盖全身自由度的扩展预训练数据,LingBot-VLA-2.0 在两个机器人平台上展现出强大的跨具身长时程移动操作能力。Benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
MV-Forcing:通过 4D 几何基础的时空自强迫实现长多视角视频生成
arXiv:2607.05376 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MV-Forcing 框架,通过在顺序生成视角之间引入 4D 几何桥梁,在单一扩散模型中组合时间与视角自回归,并弥合时间与视角序列自回归中训练-推理的曝光偏差差距。MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views, and closes the train-inference exposure bias gap for both temporal and view-sequential autoregression.

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes
AI Wizards 参加 EXIST 2026:用于迷因中多模态性别歧视识别的分层软标签学习
arXiv:2607.04410 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作通过轻量级 Gated MLP 将固定的 Gemini Embedding 2 视觉-语言表征映射到目标空间,使用 KL 散度与同方差不确定性加权进行训练,并展示了 AI Wizards 参加 EXIST 2026 多模态迷因性别歧视识别任务的提交方案。This work maps fixed Gemini Embedding 2 vision-language representations through a lightweight Gated MLP trained with KL divergence and homoscedastic uncertainty weighting, and presents the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes.

Gemma 4 Technical Report
Gemma 4 技术报告
arXiv:2607.02770 多模态 方法 OA · 绿色 被引 42 · S2

本工作推出 Gemma 4——Gemma 模型系列中新一代开源权重、原生多模态的语言模型,在 STEM、多模态与长上下文基准上实现性能跃升,在人类评分任务上可与更大的前沿开源模型相媲美。This work introduces Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family that establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
OrbitQuant:面向图像与视频扩散 Transformer 的数据无关量化
arXiv:2607.02461 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 OrbitQuant,一种数据无关的权重量化器,通过在归一化旋转基空间中进行量化以绕过范围估计,将图像扩散 Transformer 的 PTQ 推进到 W2A4 并保持可用生成质量。OrbitQuant is presented, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis and pushes PTQ of image diffusion transformers to W2A4 with usable generation quality.

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
MultAttnAttrib:长文档问答中的免训练多模态归因
arXiv:2607.01420 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 MultAttnAttrib,一种免训练的归因生成方法,利用模型的预填充过程、选定的注意力头以及校准阈值在文档中定位源证据,且在多种归因生成方法上一致地表现更优。This work introduces MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a document, and consistently outperforms a variety of attribution-generation methods.

Generated Contents Enrichment
生成内容增强
arXiv:2405.03650 多模态 方法 OA · 绿色 被引 1 · S2

本文提出一个联合训练的对抗框架,通过建模对象语义和对象间关系来增强场景图,并在 Visual Genome 数据集上以代理场景图增强指标、图像质量比较、定性示例与用户研究进行评估。A jointly trained adversarial framework is proposed that enriches scene graphs by modeling object semantics and inter-object relations and is evaluated with proxy scene graph enrichment metrics, image-quality comparisons, qualitative examples, and user studies on the Visual Genome dataset.

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
京东 Oxygen AI 商品中心(Oxygen AIIC)V1:以 LLM/VLM 为核心的工业级商品理解、管理与应用解决方案
arXiv:2606.28070 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

京东 Oxygen AI 商品中心(Oxygen AIIC):基于 LLM/VLM 的工业级商品知识生产与服务平台,已在大规模场景下取得可量化的收益。The JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service, has delivered measurable gains at scale.

DataComp-VLM: Improved Open Datasets for Vision-Language Models
DataComp-VLM:面向视觉-语言模型的改进开源数据集
arXiv:2606.28551 多模态 评测集 MPG.PuRe (Max Planck Society) OA · 绿色 被引 1 · S2

数据混合(而非过滤)是构建高质量训练数据集的关键:以指令型数据为主的混合在扩展时优于以描述型数据为主的混合,且规模越大优势越明显。It is found that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales.

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
VLA-Corrector:面向自适应动作时域的轻量级检测-修正推理
arXiv:2607.01804 多模态 方法 OA · 绿色 被引 3 · S2

VLA-Corrector:面向动作分块 VLA 策略的轻量级修正推理框架;引入轻量的潜空间视觉监控器,持续比对预测与实际视觉特征演化,可在线检测视觉动态偏差,缓解静态时域在执行鲁棒性与策略调用频率之间的权衡。VLA-Corrector, a lightweight corrective inference framework for action-chunked VLA policies, introduces a lightweight Latent-space Vision Monitor that continuously compares predicted and actual visual feature evolution, enabling online detection of visual dynamics deviations and mitigates the trade-off imposed by static horizons between execution robustness and policy-call frequency.

From SRA to Self-Flow: Data Augmentation or Self-Supervision?
从 SRA 到 Self-Flow:数据增强还是自监督?
arXiv:2607.02508 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Attention Separation,在保留与 Self-Flow 相同的双时间步输入的同时,阻止被分配到不同噪声水平 token 之间的注意力,并表明 Attention Separation 本身通过将单张图像拆分为多个有效训练部分来扩充训练数据,从而带来增强效果。Attention Separation is introduced, which preserves the same dual-timestep input as Self-Flow while blocking attention between tokens assigned to different noise levels, and shows that Attention Separation itself provides an augmentation effect by splitting a single image into multiple effective training parts to expand the training data.

Discrete Diffusion Language Models for Interactive Radiology Report Drafting
用于交互式放射学报告起草的离散扩散语言模型
arXiv:2607.01436 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文适配了一款专家混合扩散语言模型 DiffusionGemma-26B,并在医学视觉问答数据集上,使用相同的 LoRA 配置将其与同规模的自回归模型 Gemma-4-26B 进行基准对比,由对冗长度鲁棒的 LLM 裁判打分。This work adapts a mixture-of-experts diffusion language model, DiffusionGemma-26B, and benchmark it against its same-size AR sibling Gemma-4-26B under an identical LoRA recipe on medical visual question answering datasets, scored by a verbosity-robust LLM judge.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
基于非对称互变分学习的多模态连续推理
arXiv:2607.00461 多模态 方法 OA · 绿色 被引 1 · S2

AMVL 在潜空间集成的 MLLM 中实例化,持续优于多种强离散和潜空间推理基线,在复杂 BLINK 基准上平均得分提升 +10.83,单个推理任务最高提升 +32.00,分析也确认了潜空间稳定性的改善。AMVL is instantiate in a latent-integrated MLLM and it consistently outperforms strong discrete and latent-reasoning baselines, improving the average score on the complex BLINK benchmark by +10.83 and achieving gains of up to +32.00 on individual reasoning tasks, with analyses confirming improved latent-space stability.

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
打破失败级联:面向医学多模态推理的步骤感知强化学习
arXiv:2606.31825 多模态 评测集 OA · 绿色 被引 1 · S2

在三种多模态 LLM 主干模型上,MRPO 均稳定优于标准 GRPO 及一项最新的 RL 基线;在 Qwen3-VL-8B-Instruct 上甚至超越规模显著更大的医学 MLLM(如 HuatuoGPT-Vision-34B)2.79 分。Across three multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3-VL-8B-Instruct even surpasses substantially larger medical MLLMs such as HuatuoGPT-Vision-34B by 2.79 points.

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation
DataEvolver:面向富文本图像生成的自演化多 Agent 数据构建
arXiv:2606.31537 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在富文本图像生成基准上的实验表明,在数据预算匹配条件下,DataEvolver 比固定数据集基线生成更有用的训练数据,且结果表明被拒绝的样本可为改进富文本图像数据构建提供可操作的反馈信号。Experiments on text-rich image generation benchmarks show that DataEvolver produces more useful training data than fixed-dataset baselines under matched data budgets, and results suggest that rejected samples can provide actionable feedback for improving text-rich image data construction.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
QVal:低成本评估面向长 horizon LLM Agent 的密集监督信号
arXiv:2606.32034 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 QVal,一个无需训练即可直接评估密集监督信号的测试平台,发现简单提示基线始终优于文献中近期提出的密集监督方法,且性能按模型家族高度聚类。QVal, a training-free testbed for directly evaluating dense supervision signals, is introduced, finding that simple prompting baselines consistently outperform recent dense supervision methods from the literature, and that performance clusters strongly by family.

MuSViT: A Foundation Vision Model for Sheet Music Representation
MuSViT:面向乐谱表示的基础视觉模型
arXiv:2606.31811 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 MuSViT (Music Score Vision Transformer):首个面向乐谱表征的基础视觉模型——一个通过 Masked Autoencoders 在 IMSLP 970 万页数据上预训练的 ViT 编码器。This work introduces MuSViT (Music Score Vision Transformer): the first foundation vision model for sheet music representation -- a ViT encoder pre-trained via Masked Autoencoders on 9.7 million pages from the IMSLP.

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
照亮统一多模态模型:面向自由形式交错图文生成
arXiv:2606.30054 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ILLUME-X,一种先进的统一多模态范式,通过提升多模态数据效率并稳定多模态训练过程,实现高质量、自由形式的交错图文生成。This paper introduces ILLUME-X, an advanced unified multimodal paradigm that enables high-quality, free-form interleaved text-image generation by improving multimodal data efficiency and stabilizing the multimodal training process.

Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
用于流量矩阵预测的参数高效量子启发快速权重编程器
arXiv:2606.27821 多模态 方法 OA · 绿色 被引 3 · S2

本文将门控量子启发的 Kolmogorov-Arnold 网络快权重编程器用于直接多步 Abilene 流量矩阵预测,提出以经典慢速编程器搭配量子启发快速编程器的方案,作为面向资源受限场景的网络流量矩阵预测中一种兼顾精度与效率的有前景设计。This paper adapts gated quantum-inspired Kolmogorov-Arnold network fast-weight programmers to direct multi-step Abilene TM forecasting and identifies a classical slow programmer with a quantum-inspired fast programmer as a promising accuracy-efficiency design for resource-conscious network traffic-matrix forecasting.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
PhysRAG:通过检索增强生成提升视频生成中的物理感知
arXiv:2606.26916 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 PhysRAG——一条通过检索增强生成(RAG)提升视频生成物理感知的新流程,并基于 WISA-80K 数据集设计了两阶段数据过滤流程,最终筛选出 7K 高质量视频用于训练。This work introduces PhysRAG, a novel pipeline that enhances physical awareness in video generation through Retrieval-Augmented Generation (RAG), and designs a two-stage data filtering pipeline based on the WISA-80K dataset, resulting in a curated set of 7K high-quality videos for training.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages
RedVox:语音模型跨语言的安全性与公平性差距
arXiv:2606.26968 多模态 综述 OA · 绿色 被引 0 · S2 + OpenAlex

提出 RedVox,一个基于真实人声构建的音频与语音多语言安全性与公平性基准,涵盖五种语言中的不安全与不公平的刻板请求;研究发现漏洞即使在非对抗条件下仍然存在,在非英语语言中更为严重,且在请求来自语音输入时会被进一步放大。RedVox is introduced, a multilingual safety and fairness benchmark for audio and speech built on real voices, covering unsafe and unfair stereotypical requests across five languages, finding that vulnerabilities persist even under non-adversarial conditions, worsen in non-English languages, and are amplified when the request comes from a spoken input.

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
Learning to Fold:LeHome Challenge 2026 获奖方案(线上第 1,线下第 2)
arXiv:2606.27163 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作通过强化学习循环改进视觉-语言-动作(VLA)策略,该循环预测成功、进展及若干任务相关的未来量,并驱动优势估计、实时失败检测与候选选择,在 LeHome Challenge 2026 中取得佳绩。The work improves a vision-language-action (VLA) policy with a reinforcement-learning loop that predicts success, progress, and a few task-relevant future quantities and drives advantage estimation, live failure detection, and candidate selection in the LeHome Challenge 2026.

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
EO-WM:面向概率性地球观测预报的物理信息世界模型
arXiv:2606.27277 多模态 方法 OA · 绿色 被引 1 · S2

地球观测(EO)预报旨在依据变化的天气条件,从卫星观测预测未来地表动态。本文将其建模为部分可观测、天气驱动的世界建模问题,其中天气作为条件信号,而由于观测稀疏和未观测的陆面状态,预报本身具有不确定性。然而现有方法未能完整刻画这一设定:确定性模型将不确定性坍缩为单一未来预测,而基于扩散的方法通常将天气变量视作无条(原文此句截断)。Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as un

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
Dziri Voicebot:面向阿尔及利亚方言的端到端低资源语音对话系统
arXiv:2606.26003 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

所提系统为阿尔及利亚方言的端到端对话建模提供了可复现基线;实验结果显示各组件均表现优异:ASR 词错误率低,NLU 意图分类与实体识别得分高,语音合成质量稳定。The proposed system provides a reproducible baseline for end-to-end conversational modeling in Algerian Dialect, and experimental results show strong performance across all components, including low word error rate for ASR, high intent classification and entity recognition scores for NLU, and stable speech synthesis quality.

Distribution-Aware Diffusion-LLM for Robust Ultra-Long-Term Time Series Forecasting
用于鲁棒超长期时间序列预测的分布感知扩散 LLM
arXiv:2606.23391 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出新框架 Diffusion-LLM,将条件扩散模型集成到基于 LLM 的预测流水线中,展示了分布感知正则化在提升时间序列 LLM 的鲁棒性与泛化能力方面的价值。This work proposes a new framework Diffusion-LLM that integrates a conditional diffusion model into an LLM-based forecasting pipeline, and demonstrates the value of distribution-aware regularization for enhancing robustness and generalization in time series LLMs.

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
RaysUp:基于几何感知光线表示的超轻量通用特征上采样
arXiv:2606.22749 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 RaysUp,一个超轻量级、任务无关且与 VFM 无关的特征上采样框架,可在任意分辨率下重建高分辨率特征图,仅使用 AnyUp 16% 的参数即达到 SOTA 性能,推理速度提升约 7 倍。RaysUp is proposed, an ultra-lightweight, task-agnostic, and VFM-agnostic feature upsampling framework that reconstructs high-resolution feature maps at arbitrary resolutions that achieves state-of-the-art performance while using only 16% of the parameters of AnyUp and delivering approximately 7x faster inference.

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
ABACUS:适配统一基础模型以桥接图像计数理解与生成
arXiv:2606.23835 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

ABACUS 是一个统一视觉语言模型,可在无需任何基准特定训练的情况下处理物体计数、人群计数、指代表达计数以及忠实计数的图像生成,性能超越任务特定的专家模型和更大的通用模型。ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required, outperforming both task-specific specialists and larger generalist models.

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization
MedRLM:用于长上下文临床推理、传感器引导筛查、循证决策支持和社区到三级转诊优化的递归多模态健康智能
arXiv:2606.20164 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

MedRLM 旨在将医疗 AI 从静态问答转向可审计、多模态且工作流感知的临床决策支持,并引入临床证据图记忆,将患者特定观察与检索到的证据相连接。MedRLM aims to move medical AI from static question answering toward auditable, multimodal, and workflow-aware clinical decision support, and introduces a Clinical Evidence Graph Memory to connect patient-specific observations with retrieved evidence.

Vesta: A Generalist Embodied Reasoning Model
Vesta:通用具身推理模型
arXiv:2606.20905 多模态 应用落地 OA · 绿色 被引 2 · S2

提出 Vesta,一个统一的具身通用模型,将定位、空间推理、导航和长程规划能力整合到单个基础模型中,并证明通用模型能够达到或超越专家模型。Vesta is presented, a unified embodied generalist that consolidates localization, spatial reasoning, navigation, navigation, and long-horizon planning capabilities into a single foundation model and demonstrates that a generalist model can match or exceed specialists.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models
CogniRoute:全模态模型中的社交证据路由学习
arXiv:2606.20970 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CogniRoute,一种面向社交全模态推理的 schema 引导 Mixture-of-Experts(混合专家)框架,并引入路由感知强化学习,通过答案正确性、模态一致性推理与认知时序锚定等奖励联合优化 token 生成与专家分配。CogniRoute, a schema-guided Mixture-of-Experts framework for social omni reasoning, is introduced and route-aware reinforcement learning is introduced, which jointly optimizes token generation and expert allocation using rewards for answer correctness, modality-consistent reasoning, and cognitive temporal grounding.

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
基于物体中心的残差强化学习用于 VLA 零样本仿真到现实迁移的增强
arXiv:2606.18953 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种基于物体中心的残差强化学习框架,利用物体位姿精化 VLA 动作,使观测空间紧凑,在仿真与现实之间能够一致迁移。An object-centric residual RL framework is proposed that refines VLA actions using object poses, enabling a compact observation space that transfers consistently between simulation and reality.

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
RoboTALES:通过任务对齐的模拟未来学习推理引导的机器人策略
arXiv:2607.06018 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 RoboTALES,一个学习任务对齐模拟未来并据此训练机器人策略的单阶段框架,引入两项关键创新:基于 LLM 的分层规划器,将复杂任务拆解为子目标序列以引导模型的"想象";基于 VLM 的评判器,用于评估这些"想象"出的未来。This work proposes RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies and introduces two key innovations: a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination and a VLM-based critic that evaluates these ``imagined'' futures.

Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
基于 Token 的双视图融合与适配用于乳腺癌分类的大视觉模型
arXiv:2607.06309 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种以 Token 为中心的双视图学习框架,在冻结的视觉 Transformer 主干中统一基于 prompt 的适配与跨视图融合,相比线性探测、仅 prompt 适配以及传统融合基线均取得稳定提升。A token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone and demonstrates consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines.

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
为触觉而醒!MLLM 中基于掩码隔离的触觉对齐学习
arXiv:2607.00302 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

[摘要] Splash 是一个面向 MLLMs 的掩码隔离触觉对齐学习框架,它量化每个预训练参数的重要性,并将参数空间划分为休眠子空间与关键子空间,从而有效防止灾难性遗忘,确保非破坏性的模态扩展。Splash is presented, a mask-isolated tactile alignment learning framework for MLLMs that quantifies the significance of each pretrained parameter, and partitions the parameter space into a dormant and critical subspace, which effectively prevents catastrophic forgetting and ensures non-destructive modality expansion.

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
CineMobile:面向电影级相机运动生成的端侧图像到视频扩散模型
arXiv:2607.03803 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

CineMobile 采用三重优化策略,通过蒸馏引导的剪枝方法得到一个紧凑而高效的模型,保留实现电影级效果所需的核心视频生成能力,证明了其在移动端图像到视频创作中的实用可行性。CineMobile adopts a three-fold optimization strategy, leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects, demonstrating its practical applicability for mobile-based image-to-video creation.

Video-Oasis: Rethinking Evaluation of Video Understanding
Video-Oasis:重新审视视频理解评估
arXiv:2603.29616 多模态 评测集 OA · 绿色 被引 2 · S2

该审计揭示现有基准样本中 55% 可在无视觉输入或时序上下文的情况下被解决,并提出 Video-Oasis,一个用于系统性审计现有视频理解基准的可持续诊断套件。This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context, and introduces Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks.