提出 LingBot-Video,一种专为具身智能设计的基于 DiT 的视频预训练范式,并将其作为首个大规模开源 MoE 视频基础模型贡献给社区,致力于在数字创意与物理执行之间搭建桥梁。LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.
论文
417 张论文卡片 · 多模态 · OA 绿色
提出了 MotorMind,一个机器人操作框架,将 VLM 提出的中级动作与确定性机器人控制及反馈相连接,支持异步监控与后台记忆更新;研究表明,通用 VLM 在配备合适的中级动作表示和异步执行框架后,可执行有效的零样本机器人操作。MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.
提出 ProWAM,一种渐进式世界动作模型,联合预测动作和稀疏视觉子目标的有序序列,为任务执行过程中的动作生成提供显式视觉引导,展示了进度索引视觉前瞻对闭环控制的价值。ProWAM is presented, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution, demonstrating the value of progress-indexed visual foresight for closed-loop control.
提出 TQP,一种训练策略,将视频按帧分块处理,块间传递被追踪对象信息,但仅在单个块内进行反向传播,从而在无显存溢出、推理开销或梯度消失问题的情况下实现更长时间范围的监督。TQP is introduced, a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients.
提出 Rollout-Marginal Distillation,在 AR 预测中保留生成历史,但对每个 chunk 独立对照 chunk teacher 打分,确保其质量修正不受不完美的时序上下文影响。Rollout-Marginal Distillation is introduced, which retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context.
一种 3D 世界动作模型,将世界分解为场景与手,并在共享时空坐标系内将二者联合预测为 3D 点轨迹,无需任务特定物体或关键点选择即可在大量人类演示视频上进行有效预训练。A 3D world action model that decomposes the world into a scene and hands and hands, and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame, which enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection.
EyeRobot 2.0 通过旋转两个眼动视角将其注视中心对准场景中的 3D 注视点实现物理上的视觉聚焦,并将夹爪信息规范化到以注视点为参考的 SE(3) 坐标系,从而压缩待学习的动作分布规模。EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it, and takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn.
首次在 World Modeling 领域引入 Agentic Harness 框架:由 pilot agent 负责规划并执行角色行为,director agent 负责随着场景推进合成新的环境要素。The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.
提出 WM-VLM,为预训练 VLM 配备一个轻量级世界模型分支以生成中间视觉状态,表明内部世界模型为 VLM 在语言与视觉双重空间中的推理提供了一条有前景的路径。WM-VLM is introduced, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states and suggests that internal world models offer a promising path toward VLMs that reason in both language and visual space.
提出 ACG-Bench,一个面向逐臂组合泛化(arm-wise Compositional Generalization)的基准,为超越固定训练流程的双臂策略设计提供实证指导,考察了 arm-token grouping、技能专属 LoRA 适配器(SkillLoRA)以及 arm-wise attention(AWA),凸显了技能条件化参数与注意力结构的互补性。This work introduces ACG-Bench, a benchmark for arm-wise Compositional Generalization that provides empirical guidance for designing dual-arm policies that generalize beyond fixed training routines, and examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure.
提出 LaMem-VLA,一种以潜在记忆为核心框架的方法,将历史经验重建为潜在记忆 token,并直接与 VLA 推理交织,使记忆能够在有界上下文下直接参与 VLA 推理并引导动作生成。LaMem-VLA is introduced, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning, and enables memory to directly participate in VLA reasoning and guide action generation under a bounded context.
Perturbot 和 GroundingFscore 提供了一套训练与评估框架,用于将 VLA 决策与捷径先验解耦,同时保持对任务相关证据的响应性。Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.
论文提出 Self-compensating VLA,一种部署阶段的自适应方法,使 VLA 策略在生成指令时能够预补偿机器人的执行误差,并在平均任务成功率上高于基线策略以及在训练阶段增强鲁棒性的方法。Self-compensating VLA is proposed, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands, and achieves higher average task success than both the base policies and methods that build in robustness during training.
本文提出 DiffGate,一种将 GRPO 与选择性、有界教师指导相结合的结果门控目标,在四种模型–领域设置下均提升了 pass@8,表明在该评估协议下解的覆盖度得到改善。DiffGate is introduced, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance, and improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under the evaluation protocol.
自回归视频扩散支持流式生成与交互控制,但其 KV cache 随生成历史持续增长。现有压缩策略要么采用固定窗口丢弃历史,要么依据局部注意力与相似度信号选择 token,均无法直接衡量当前块是否贡献了超出已保留上下文的信息。本文提出 DeCoPrune,一种将 cache 压缩视为去噪一致性问题的免训练方法。经验上发现,去噪难度可作为 token 价值的有用代理……Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's valu
分布匹配蒸馏(DMD)通过分别估计的目标分数与学生分数之差训练少步学生模型,因而必须维护一个拟合学生动态分布的辅助扩散模型,带来额外的显存与计算开销。本文提出 DMAD——分布匹配对抗蒸馏,将分布匹配重构为分类任务,直接学习所需的 log-density 比。共享主干上的两个判别头分别区分真实数据与教师样本是否来自学生模型,并对其 logits 施加线性损失以训练……Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the
将音乐转录为人类可读的乐谱需要对节奏、和声、旋律与曲式的整体理解。两大障碍限制了这一目标的实现:标注录音稀缺,以及精确的局部预测仍可能产生不一致的音乐序列。本文提出 SheetSage2,一个统一的音乐转录框架,结合合成数据、任务特定的结构化解码与自回归蒸馏。自动标注的 MIDI 渲染为音频后,可为各类音乐理解任务提供可扩展的监督。任务特定的结构化解码器整合互补的音乐……Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary mu
稀疏视图新视角合成是三维内容创作中的核心问题,但基于扩散的方法受迭代去噪限制,多视图生成在推理时开销高昂。本文提出 NAMVIS,一种无扩散框架,将多视图图像合成重新表述为几何条件下的下一尺度自回归过程。NAMVIS 不通过反复去噪生成目标视图,而是通过少量由粗到细的尺度步预测离散视觉 token,并在同一尺度内以及跨目标视图间并行采样所有 token。为锚定此过程……Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor thi
视觉-语言-动作(VLA)基础模型规模迅速扩大以提升操作性能与泛化能力,但这种规模化带来了高昂的计算成本,使真实世界部署日益困难。现有方法通常通过设计更小的架构或减少基于流(flow-based)策略中的迭代去噪步数来缓解该问题。本文提出 FastOPD,一个从基础到轻量的 VLA 框架,通过高效的在线蒸馏实现大规模 VLA 的实际部署。具体而言,FastOPD 适配流映射(flow map)……Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map
WildCity 是一个由自动驾驶车队在复杂城市环境中采集的真实多模态数据集,旨在推动城市级渲染的进展,并更广泛地推动 AI 在空间感知、记忆与推理方面达到与人类认知相当规模的能力。WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments, aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition.
得益于涵盖全身自由度的扩展预训练数据,LingBot-VLA-2.0 在两个机器人平台上展现出强大的跨具身长时程移动操作能力。Benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.
提出 MV-Forcing 框架,通过在顺序生成视角之间引入 4D 几何桥梁,在单一扩散模型中组合时间与视角自回归,并弥合时间与视角序列自回归中训练-推理的曝光偏差差距。MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views, and closes the train-inference exposure bias gap for both temporal and view-sequential autoregression.
该工作通过轻量级 Gated MLP 将固定的 Gemini Embedding 2 视觉-语言表征映射到目标空间,使用 KL 散度与同方差不确定性加权进行训练,并展示了 AI Wizards 参加 EXIST 2026 多模态迷因性别歧视识别任务的提交方案。This work maps fixed Gemini Embedding 2 vision-language representations through a lightweight Gated MLP trained with KL divergence and homoscedastic uncertainty weighting, and presents the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes.
本工作推出 Gemma 4——Gemma 模型系列中新一代开源权重、原生多模态的语言模型,在 STEM、多模态与长上下文基准上实现性能跃升,在人类评分任务上可与更大的前沿开源模型相媲美。This work introduces Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family that establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.
本文提出 OrbitQuant,一种数据无关的权重量化器,通过在归一化旋转基空间中进行量化以绕过范围估计,将图像扩散 Transformer 的 PTQ 推进到 W2A4 并保持可用生成质量。OrbitQuant is presented, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis and pushes PTQ of image diffusion transformers to W2A4 with usable generation quality.
本工作提出 MultAttnAttrib,一种免训练的归因生成方法,利用模型的预填充过程、选定的注意力头以及校准阈值在文档中定位源证据,且在多种归因生成方法上一致地表现更优。This work introduces MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a document, and consistently outperforms a variety of attribution-generation methods.
本文提出一个联合训练的对抗框架,通过建模对象语义和对象间关系来增强场景图,并在 Visual Genome 数据集上以代理场景图增强指标、图像质量比较、定性示例与用户研究进行评估。A jointly trained adversarial framework is proposed that enriches scene graphs by modeling object semantics and inter-object relations and is evaluated with proxy scene graph enrichment metrics, image-quality comparisons, qualitative examples, and user studies on the Visual Genome dataset.
京东 Oxygen AI 商品中心(Oxygen AIIC):基于 LLM/VLM 的工业级商品知识生产与服务平台,已在大规模场景下取得可量化的收益。The JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service, has delivered measurable gains at scale.
数据混合(而非过滤)是构建高质量训练数据集的关键:以指令型数据为主的混合在扩展时优于以描述型数据为主的混合,且规模越大优势越明显。It is found that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales.
VLA-Corrector:面向动作分块 VLA 策略的轻量级修正推理框架;引入轻量的潜空间视觉监控器,持续比对预测与实际视觉特征演化,可在线检测视觉动态偏差,缓解静态时域在执行鲁棒性与策略调用频率之间的权衡。VLA-Corrector, a lightweight corrective inference framework for action-chunked VLA policies, introduces a lightweight Latent-space Vision Monitor that continuously compares predicted and actual visual feature evolution, enabling online detection of visual dynamics deviations and mitigates the trade-off imposed by static horizons between execution robustness and policy-call frequency.
提出 Attention Separation,在保留与 Self-Flow 相同的双时间步输入的同时,阻止被分配到不同噪声水平 token 之间的注意力,并表明 Attention Separation 本身通过将单张图像拆分为多个有效训练部分来扩充训练数据,从而带来增强效果。Attention Separation is introduced, which preserves the same dual-timestep input as Self-Flow while blocking attention between tokens assigned to different noise levels, and shows that Attention Separation itself provides an augmentation effect by splitting a single image into multiple effective training parts to expand the training data.
本文适配了一款专家混合扩散语言模型 DiffusionGemma-26B,并在医学视觉问答数据集上,使用相同的 LoRA 配置将其与同规模的自回归模型 Gemma-4-26B 进行基准对比,由对冗长度鲁棒的 LLM 裁判打分。This work adapts a mixture-of-experts diffusion language model, DiffusionGemma-26B, and benchmark it against its same-size AR sibling Gemma-4-26B under an identical LoRA recipe on medical visual question answering datasets, scored by a verbosity-robust LLM judge.
AMVL 在潜空间集成的 MLLM 中实例化,持续优于多种强离散和潜空间推理基线,在复杂 BLINK 基准上平均得分提升 +10.83,单个推理任务最高提升 +32.00,分析也确认了潜空间稳定性的改善。AMVL is instantiate in a latent-integrated MLLM and it consistently outperforms strong discrete and latent-reasoning baselines, improving the average score on the complex BLINK benchmark by +10.83 and achieving gains of up to +32.00 on individual reasoning tasks, with analyses confirming improved latent-space stability.
在四个多模态 LLM backbone 上,MRPO 始终优于标准 GRPO 与一项最新 RL baseline,并在 Qwen3-VL-8B-Thinking 上以 4.59 分超越规模远大于它的医学 MLLM(如 HuatuoGPT-Vision-34B)。Across four multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3-VL-8B-Thinking even surpasses substantially larger medical MLLMs such as HuatuoGPT-Vision-34B by 4.59 points.
在富文本图像生成基准上的实验表明,在数据预算匹配条件下,DataEvolver 比固定数据集基线生成更有用的训练数据,且结果表明被拒绝的样本可为改进富文本图像数据构建提供可操作的反馈信号。Experiments on text-rich image generation benchmarks show that DataEvolver produces more useful training data than fixed-dataset baselines under matched data budgets, and results suggest that rejected samples can provide actionable feedback for improving text-rich image data construction.
本文提出 QVal,一个无需训练即可直接评估密集监督信号的测试平台,发现简单提示基线始终优于文献中近期提出的密集监督方法,且性能按模型家族高度聚类。QVal, a training-free testbed for directly evaluating dense supervision signals, is introduced, finding that simple prompting baselines consistently outperform recent dense supervision methods from the literature, and that performance clusters strongly by family.