视频扩散模型具有随机性且难以控制:精确内容往往需要反复采样且无法保证成功,长时场景在外观、交互和时间一致性上会发生漂移。Agent式视觉创作可提供显式参考、可编辑的3D场景或可执行的游戏状态以实现稳定控制,但本身并不能保证高对象或角色保真度。二者结合可实现稳定且高质量的生成。为实现该结合,我们提出 LynnReal-Omni,一个基于32B共享多模态Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or executable game states for stable control, but does not by itself guarantee high object or character fidelity. Combining the two can enable stable, high-quality generation. To realize this combination, we present LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimo
论文
417 张论文卡片 · 多模态 · OA 绿色
ModaLens 是一种配对图像交换审计,用于衡量报告可用性如何改变图像敏感性,并限制对视觉正确性的结论;该方向在另外两个模型系列中得到复现。ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity and limits conclusions about visual correctness; the direction replicates in two further model lineages.
Dynin-Robotics 在 LIBERO 和零样本 LIBERO-Plus 上取得了具有竞争力的性能,并在 Franka Research 3 机器人的四种操作条件下达到了 78.4% 的平均成功率。Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.
PhysStream 是一个用于物理驱动图像到视频合成的自回归模型,通过引入结构化场景记忆并支持基于稀疏速度增量信号的细粒度运动控制(编码物理量),使模型能够学习底层动力学。PhysStream is an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics.
StepAudio 3 Realtime 是一个围绕持续 listen-converse-think-act 循环组织的音频-语言基础模型,在实时语音输出的同时达到与专用推理模型相当的对话与推理性能,并通过 Think-While-Speaking 机制化解深度思考与延迟之间的矛盾。StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop, achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time, and resolves the tension between deep deliberation and latency via Think-While-Speaking.
该工作提出了 StepAudio 3 Music,一个支持显式音乐规划与开放域文本控制生成的大规模长篇音乐生成模型,在所评估系统中取得最高的 AudioBox Content Enjoyment、Content Usefulness 与 Production Quality 分数,以及最高的 MuQ-MuLan 相似度。This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems.
Register被实现为专用的固定位置 token,其连续的隐藏状态被训练用于跨生成块承载推理进度,对有界代码生成尤其有效,因为正确程序通常跨越多个块。Registers are implemented as dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks, which are especially effective for bounded code generation, where correct programs usually span several chunks.
本文提出一个受控的跨模态框架,在多种模态下实例化相同的任务套件以检验 Convergent Emergence Hypothesis,结果显示配对映射的 ICL 在六种模态中出现,超越受控基线,并在其中五种模态上呈现相关的逐任务效应。A controlled cross-modality framework that instantiates the same task suite in a variety of modalities to test the Convergent Emergence Hypothesis and shows that paired-mapping ICL emerges across six modalities, surpasses controlled baselines, and has correlated per-task effects across five of them.
本文提出物理秩一致性 (PRC) 来衡量 tokenization 在重建后保留局部物理距离排序的程度,并提出 ActionPiece,通过对表示学习和量化的联合监督来保留物理动作关系。This work introduces physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction, and presents ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization.
我们提出 Zing-0.5,一个 5B 自回归世界模型,专注于可玩性:用户可以探索生成的世界、影响事件演进,并通过键盘与在线文本的联合控制对反馈做出响应。我们的方法汇聚了三项技术贡献:(1) 统一的动作与文本条件建模,将感知幅度的键盘输入与时序对齐的文本指令、以及联合标注的视频结合,在同一序列中学习导航与事件控制;(2) 面向增量生成的事件尺度监督,使用分段级教师...We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trai
本文提出 PanoCaps,一个由全景分割数据集构建的人工标注 benchmark,并提出 PANORAMA,一个将预训练 segmenter 条件化于上下文化短语表示以获得候选 mask、并学习选择与每个短语对应的 mask 的 VLM。This work introduces PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets, and introduces PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase.
提出 UFO,这是首个面向全条件对齐同时评估的统一框架;并发布 UFO-Bench,一个用于整体评估现有定制化模型在文本与视觉条件多样化交互下表现的专用基准。UFO is proposed, the first unified framework for omni-condition alignment simultaneous evaluation, and UFO-Bench is presented, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.
本文提出 FAMOS,一个前馈模型,可从稀疏、无序的部分点云集合预测可动部件分割与关节参数;并引入一个过程式数据生成器,在训练过程中合成自标注资产,以克服现有数据集规模和多样性的局限。FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds, is presented and a procedural data generator that synthesizes self-annotated assets during training is introduced to overcome the limited scale and diversity of existing datasets.
本文提出一个围绕物理世界推理四个互补维度构建的综合评估框架,并构建了一系列多样化新任务,要求模型跨模态整合互补信息。This work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning, and constructs a diverse set of novel tasks that require models to integrate complementary information across modalities.
提出 Paint-Anything,通过物体级颜色监督学习一个统一的 hex-prompt 界面以同时支持生成与编辑;并引入 Any Color Benchmark (ACBench),包含 ACBench-T2I 和 ACBench-Edit,以衡量两项任务中的物体级 hex 颜色保真度。Paint-Anything is presented, which learns a shared hex-prompt interface for generation and editing through object-level color supervision, and introduces Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks.
提出 FRAUDSkill——一种结构化的 frozen-weight 适配框架:底层音频-语言模型保持不变,转而优化外部的 skill program、路由策略和决策规则,并将结构化输出控制与验证引导的多路径推理相结合,以确保符合协议规范的预测。FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions.
HEAL 将动态信息校准因子注入协同注意力头的 value 向量中,主动调节视觉-语言依赖,引导输出分布趋向事实证据,提供了一条简单且可解释的增强模型可信度的路径。HEAL injects dynamic information calibration factors into the value vectors of synergy heads and actively regulates visual-language dependencies, steering the output distribution towards factual evidence, offering a simple and interpretable pathway to enhance model trustworthiness.
OmniVChat-Studio,一种用于合成单轮与多轮音视频对话的多 Agent 数据引擎;以及 OmniVChat-RL,一种在 OmniVChat 中同时针对回复正确性、效率与风格设计奖励的强化学习奖励方案。OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues and OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat.
提出 SteerDuplex,一种基于 Moshi 的全双工语音模型,在自然对话和针对指令遵循、语音表达、推理与双工交互的合成对话上进行了微调,并通过两阶段混合奖励强化学习改善时序与回复连贯性。SteerDuplex is introduced, a Moshi-based full-duplex speech model fine-tuned on natural conversations and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction, and two-stage reinforcement learning with hybrid rewards to improve timing and response continuity.
基于信噪分解与候选集近似的 IER(信息效率比),可基于 IER 及其与现有效用分数的组合进行 token 选择,同时保留采样的反向 KL 训练目标。An information-efficiency ratio (IER) based on a signal-to-noise decomposition and a candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective.
提出 CARE(Corrective Atomic Robotic Execution),一种通过从执行中遇到的失败进行学习以提升恢复能力的框架,并引入 Failure State Recovery Benchmark (FSR-Bench),用于评估在局部偏差与结构异常下从中间失败态恢复的表现。This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies.
提出 Grounded Action Models (GAMs),一种基于 3D grounding 构建的机器人基础模型新范式,可自主运行,并作为底层控制器由高层规划器通过多种输入模态进行控制,支持长时序与依赖记忆的操作。Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding, are proposed, which can be run autonomously and serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation.
提出 SVEET 框架,仅需在预训练的双向视频扩散模型上训练,即可支持高质量自回归流式视频编辑,并提出解耦训练方案,显式强制视频可控性与模型因果性优化方向之间的正交性。This paper proposes SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion and proposes a decoupled training scheme that explicitly enforces the orthogonality between the optimization directions of video controllability and model causality.
提出 Mira-Scene,一种组合式 3D 场景重建框架,将稀疏姿态回归替换为密集有界对应恢复,并引入多模态扩散 Transformer 联合生成物体几何与 CCM,使用模态专精的专家流配合共享注意力与位置编码以促进几何-布局一致性。Mira-Scene is presented, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery and introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency.
提出 UltraTex,一种用于高分辨率多视图扩散 3D 纹理生成的高效端到端框架,并引入 Background Token Dropping(在 DiT 主干前移除背景 token)与 Block-Sparse Attention(降低前景序列上的注意力计算)。This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.
结果表明,将部分建模负担从网络参数转移至结构化输入表示,可在更低的数据、内存与算力需求下支持紧凑的多任务视觉系统。The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.
本文构建了 TextMuSS-10M,一个涵盖 10 种文字、229 种语言的大规模合成场景文本数据集,并提出了 ScriptMoE,一种具备文字感知能力的 Mixture-of-Experts (MoE) 架构。该架构在精度上达到最高,且比 per-language experts 更简单、比 VLM 更轻量,同时精度优于两者。This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.
介绍 Tri-PvP,一个包含 8,000 样本、跨越视觉、音频与文本的三模态冲突基准,揭示了证据形式偏倚的系统性不对称:模型在视觉上更偏向感知信号,而在音频上更偏向命题性信号。Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, is introduced, revealing a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio.
介绍 Spatial-Interactor,一个通过交互训练 VLM 建模物理世界状态转换的框架,并将学习过程组织为三级课程:L1 被动世界状态转换、L2 主动自我状态转换、L3 长时交互轨迹。Spatial-Interactor is introduced, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories.
本文提出 Uranus,一个以关节轨迹为条件的自回归扩散模型为核心的数据驱动机器人仿真器,为不同机器人本体与相机配置下的同步多视图生成提供统一接口。This work presents Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations.
本文提出 X-Planner,一个面向具身推理的规划前端,同时解决监督与表征问题,并描述了规划文本质量与下游执行情况。This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.
Knowledge Pull Requests 相比从来源重写或从零重新生成,能整合更多信息并更好地保留已有内容,同时每生成一个 token 增加的信息量最多。Knowledge Pull Requests integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated.
本文提出 ViRDM,一种无需教师与评论家网络的视频后训练方案,将三网络蒸馏转化为仅生成器的后训练,在降低 GPU 显存与训练时间的同时提升视频质量。ViRDM, a teacher- and critic-free video post-training recipe that turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality, is introduced.
本文提出 OmniEchoBench,一个面向空间音视频感知与音-视-语言导航的统一基准,以及一个空间感知的全模态模型,该模型在预训练语义音频通路之外引入 FOA 空间编码器。OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation, and a spatially aware omni-modal model, which introduces an FOA spatial encoder alongside a pretrained semantic audio pathway.
结果显示 SAE 以分布式、类别依赖的形式定位形态句法信息,而非通过原子化的语法特征。The results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
本文提出 DeltaWAM,通过 dense-anchor、sparse-delta 与 action 三条流联合预测视觉 delta 与动作,并设计了三种在表征与计算共享上有所不同的架构;同时开发了 Streaming Delta Memory (SDM),使用紧凑的观测 delta 更新缓存的 anchor 上下文,从而减少繁重的 video-expert 处理。This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing.