研究库 论文知识库
Papers · organized/paper_cards

论文

443 张论文卡片 · 多模态

开放获取 全部 绿色 · 1640
PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
PreviewDiff:基于多模态评论引导的扩散潜空间搜索
arXiv:2609.36199 多模态 观点 被引 0 · S2

扩散模型能生成惊艳的图像与视频,但在组合性细节(物体数量、属性绑定、空间关系、时序动作等)上仍难以严格遵循 prompt。提升 prompt 满足度的常见做法是在测试时通过 Best-of-N 采样投入更多算力,但最终样本的选择是固定的:Best-of-N 只能从已完成的输出中挑选,无法在有潜力的轨迹失败前予以修复。我们提出 PreviewDiff,一种无需训练、基于测试时搜索的方法,将扩散采样转化为Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion samplin

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
UniEvo-VL:一种用于多模态模型自我改进的 On-policy 自蒸馏训练方案
arXiv:2609.38721 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 UniEvo-VL,一种面向多模态模型的自演化框架,可在 test-time compute 阶段从这种建设性的自纠错反馈中学习,在无外部监督或指导的情况下提升用户使用多模态模型的体验。UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute, is introduced to enhance the user experience when using multimodal models without external supervision or guidance.

Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
揭示状态之谜:Masked Diffusion Language Models 中状态适应何时重要
arXiv:2609.33355 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

转移层面的结果表明,state adaptation 选择性而非统一地应用时最为有效,且 adaptation 的价值取决于策略反转的频率与幅度。The transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly, and that adaptation value depends on both the frequency and magnitude of strategy reversals.

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
关键不在图像本身:无关上下文会扰动 VLM 评判且不提供有效信息
arXiv:2609.37863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MIST(Misleading-Image Stress Test):200 个英文句子,每句围绕一个可作比喻或字面理解的短语,配以对齐图像(描绘其读法)、误导图像(描绘相反读法)或无图像三种条件。MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.

Generative Moment Matching Networks
生成矩匹配网络
arXiv:1502.02761 多模态 方法 OA · 绿色 被引 954 · S2

本文提出一种方法,通过多层感知机的一次前馈传播生成独立样本(与近期提出的 GAN 类似),并使用 MMD 学习生成可被解码为样本的 codes。This work forms a method that generates an independent sample via a single feedforward pass through a multilayer perceptron, as in the recently proposed generative adversarial networks, using MMD to learn to generate codes that can then be decoded to produce samples.

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS:从自回归解码到掩码预测的高效免转录语音克隆
arXiv:2609.38658 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Tacit-TTS 是一个从 IndexTTS2 蒸馏而来的高效无需转录的零样本语音克隆系统,用掩码非自回归生成替换自回归的文本到语义解码,引入无需训练的声学长度估计,并通过 ReFlow 蒸馏加速流匹配渲染器。Tacit-TTS is presented, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation.

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
PixelUMM:无编码器的统一图像与视频理解与生成
arXiv:2609.38597 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding

LOCI: Spatial Linear Memory for Streaming World Models
LOCI:面向流式世界模型的空间线性记忆
arXiv:2609.40222 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LOCI,一种混合的空间记忆架构,同时维护键值缓存与循环记忆两种表示,在重访内容的复现上比代表性世界模型以及同配置的 full-softmax 模型都更为忠实。LOCI is introduced, a hybrid spatial-memory architecture that keeps both representations of key-value caches and recurrent memory that reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.

Generative Adversarial Text to Image Synthesis
生成对抗式文本到图像合成
arXiv:1605.05396 多模态 方法 OA · 绿色 被引 3461 · S2

提出一种新颖的深度架构与 GAN 形式化方法,有效衔接文本与图像建模领域的进展,将视觉概念从字符转换为像素。A novel deep architecture and GAN formulation is developed to effectively bridge advances in text and image modeling, translating visual concepts from characters to pixels.

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Ego2Act:评估自我中心视频生成中的目标导向操作
arXiv:2610.01092 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Ego2Act:一个面向目标的基准,包含来自 110 个真实日常任务的 2.640 段视频,覆盖不同的物体杂乱度与多步复杂度;同时提出 Ego2ActJudge,一种无参考评估流水线,在任务完成度与物理合理性评估上与人类共识的对齐效果优于相关基线。Ego2Act is introduced, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity, and Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines.

Video Generation Models: A Survey of Post-Training and Alignment
视频生成模型:后训练与对齐综述
arXiv:2610.00812 多模态 综述 被引 3 · S2

这篇综述旨在为推进可控且可靠的视频生成模型及帧后训练提供统一的框架,并基于对齐信号的施加方式区分隐式对齐与显式对齐,从而给出结构化的概念基础与实践指导。This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models and frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced.

Honeycomb: Constant-Size Scene Memory Representation for Video World Models
Honeycomb:面向视频世界模型的恒定大小场景记忆表征
arXiv:2609.37690 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 Honeycomb,一种基于 HexMemory 构建的视频世界模型;HexMemory 是作者提出的低秩表示,用于在固定大小的内存中存储场景特征,共包含六个空间与时空平面。Honeycomb is introduced, a video world model built on HexMemory, the authors' proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes.

World Embedding Benchmark
World Embedding Benchmark(世界嵌入基准)
arXiv:2610.03632 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 World Embedding Benchmark,包含来自 80 个族的 8,000 个受控仿真案例,涵盖流体力学、固体力学、动力学以及光学与电磁学,旨在强调需要联合评估物理一致性与物理属性可恢复性,并展示物理表示对改进视频生成的效用。The World Embedding Benchmark is introduced, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics&electromagnetism, to highlight the need to evaluate physical alignment and property recoverability jointly and demonstrate the utility of physical representations for improving video generation.

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
HyperBrowseComp:面向网页浏览 Agent 的多语言多模态压力测试
arXiv:2610.03574 多模态 评测集 被引 0 · S2

提出了 HyperBrowseComp,一个多语言、多模态的浏览基准,包含由母语或高水平使用者编写并经人工验证的、跨 13 种语言的 423 个问题,难度源于在开放网络上发现并关联证据。HyperBrowseComp is introduced, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers, with difficulty arising from discovering and connecting evidence on the open web.

ProAR: Learning Prospective Reasoning with Autoregressive Video Models
ProAR:基于自回归视频模型的前瞻性推理学习
arXiv:2610.03664 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ProAR 引入两个关键组件:一是将生成锚定到长远结果,二是通过非对称注意力掩码将目标帧预测集成到自回归循环中,使预测的目标帧可指导中间状态的生成而不被其干扰。ProAR introduces two key components: to anchor generation to the long-range outcome, and to integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
面向具身智能的 Mixture-of-Experts 视频预训练规模化
arXiv:2607.07675 多模态 方法 OA · 绿色 被引 15 · S2

提出 LingBot-Video,一种专为具身智能设计的基于 DiT 的视频预训练范式,并将其作为首个大规模开源 MoE 视频基础模型贡献给社区,致力于在数字创意与物理执行之间搭建桥梁。LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
MotorMind:为通用视觉语言模型搭建零样本机器人操控脚手架
arXiv:2609.38078 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 MotorMind,一个机器人操作框架,将 VLM 提出的中级动作与确定性机器人控制及反馈相连接,支持异步监控与后台记忆更新;研究表明,通用 VLM 在配备合适的中级动作表示和异步执行框架后,可执行有效的零样本机器人操作。MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.

World Action Modeling with Progressive Visual Planning
World Action Modeling with Progressive Visual Planning:基于渐进视觉规划的世界动作建模
arXiv:2610.02508 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ProWAM,一种渐进式世界动作模型,联合预测动作和稀疏视觉子目标的有序序列,为任务执行过程中的动作生成提供显式视觉引导,展示了进度索引视觉前瞻对闭环控制的价值。ProWAM is presented, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution, demonstrating the value of progress-indexed visual foresight for closed-loop control.

LVMT: Video Mask Transformer for Long-term Video Segmentation
LVMT:用于长期视频分割的视频掩码 Transformer
arXiv:2609.34895 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 TQP,一种训练策略,将视频按帧分块处理,块间传递被追踪对象信息,但仅在单个块内进行反向传播,从而在无显存溢出、推理开销或梯度消失问题的情况下实现更长时间范围的监督。TQP is introduced, a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients.

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation:面向长程自回归视频生成的 Rollout-Marginal 蒸馏
arXiv:2609.37925 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Rollout-Marginal Distillation,在 AR 预测中保留生成历史,但对每个 chunk 独立对照 chunk teacher 打分,确保其质量修正不受不完美的时序上下文影响。Rollout-Marginal Distillation is introduced, which retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context.

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation
PointWAM:面向灵巧机器人操作的 3D World Action Modeling
arXiv:2610.02840 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种 3D 世界动作模型,将世界分解为场景与手,并在共享时空坐标系内将二者联合预测为 3D 点轨迹,无需任务特定物体或关键点选择即可在大量人类演示视频上进行有效预训练。A 3D world action model that decomposes the world into a scene and hands and hands, and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame, which enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection.

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
EyeRobot 2.0:无需腕部摄像头的精确操作主动注视
arXiv:2610.03710 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

EyeRobot 2.0 通过旋转两个眼动视角将其注视中心对准场景中的 3D 注视点实现物理上的视觉聚焦,并将夹爪信息规范化到以注视点为参考的 SE(3) 坐标系,从而压缩待学习的动作分布规模。EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it, and takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn.

Infinite Worlds with Versatile Interactions
支持多样化交互的无限世界
arXiv:2607.07534 多模态 方法 OA · 绿色 被引 39 · S2

首次在 World Modeling 领域引入 Agentic Harness 框架:由 pilot agent 负责规划并执行角色行为,director agent 负责随着场景推进合成新的环境要素。The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
WM-VLM:探查用于交错式视觉-文本推理的内部世界模型
arXiv:2609.34826 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 WM-VLM,为预训练 VLM 配备一个轻量级世界模型分支以生成中间视觉状态,表明内部世界模型为 VLM 在语言与视觉双重空间中的推理提供了一条有前景的路径。WM-VLM is introduced, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states and suggests that internal world models offer a promising path toward VLMs that reason in both language and visual space.

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models
双臂视觉-语言-动作模型中的按臂组合泛化
arXiv:2610.06184 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ACG-Bench,一个面向逐臂组合泛化(arm-wise Compositional Generalization)的基准,为超越固定训练流程的双臂策略设计提供实证指导,考察了 arm-token grouping、技能专属 LoRA 适配器(SkillLoRA)以及 arm-wise attention(AWA),凸显了技能条件化参数与注意力结构的互补性。This work introduces ACG-Bench, a benchmark for arm-wise Compositional Generalization that provides empirical guidance for designing dual-arm policies that generalize beyond fixed training routines, and examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
面向机器人操作的视觉-语言-动作模型中的双潜在记忆
arXiv:2607.07608 多模态 方法 OA · 绿色 被引 3 · S2

提出 LaMem-VLA,一种以潜在记忆为核心框架的方法,将历史经验重建为潜在记忆 token,并直接与 VLA 推理交织,使记忆能够在有界上下文下直接参与 VLA 推理并引导动作生成。LaMem-VLA is introduced, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning, and enables memory to directly participate in VLA reasoning and guide action generation under a bounded context.

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training
PerturBot: 通过扰动训练打破视觉-语言-动作模型的捷径先验
arXiv:2610.04616 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Perturbot 和 GroundingFscore 提供了一套训练与评估框架,用于将 VLA 决策与捷径先验解耦,同时保持对任务相关证据的响应性。Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing
在机器人执行误差下驯服 VLA:自补偿与压力测试
arXiv:2609.37334 多模态 应用落地 OA · 绿色 被引 1 · S2

论文提出 Self-compensating VLA,一种部署阶段的自适应方法,使 VLA 策略在生成指令时能够预补偿机器人的执行误差,并在平均任务成功率上高于基线策略以及在训练阶段增强鲁棒性的方法。Self-compensating VLA is proposed, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands, and achieves higher average task success than both the base policies and methods that build in robustness during training.

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation
DiffGate:面向 On-Policy 蒸馏的难度门控教师指导
arXiv:2610.04596 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DiffGate,一种将 GRPO 与选择性、有界教师指导相结合的结果门控目标,在四种模型–领域设置下均提升了 pass@8,表明在该评估协议下解的覆盖度得到改善。DiffGate is introduced, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance, and improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under the evaluation protocol.

Sensor-Language-Action Models
Sensor-Language-Action 模型(感知-语言-动作模型)
arXiv:2610.08244 多模态 方法

传感器不仅有助于理解世界,还有助于决定下一步动作。然而现有传感器模型大多止于感知层面:识别状态或预测结果,而动作则通过任务特定且通常封闭的标签空间单独建模。本文提出感知-语言-动作(SLA)建模,一个在统一模型中连接多模态传感器观测、自然语言与动作的框架。SLA 将语言用作感知与动作之间的语义接口,使异构动作得以表示、预测与解释……Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while

DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
DeCoPrune:通过去噪一致性实现自回归视频扩散的高效 KV-Cache 剪枝
arXiv:2609.39096 多模态 方法 OA · 绿色 被引 0 · OpenAlex

自回归视频扩散支持流式生成与交互控制,但其 KV cache 随生成历史持续增长。现有压缩策略要么采用固定窗口丢弃历史,要么依据局部注意力与相似度信号选择 token,均无法直接衡量当前块是否贡献了超出已保留上下文的信息。本文提出 DeCoPrune,一种将 cache 压缩视为去噪一致性问题的免训练方法。经验上发现,去噪难度可作为 token 价值的有用代理……Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's valu

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
DMAD:面向快速视觉生成的分布匹配对抗蒸馏
arXiv:2610.02188 多模态 方法 OA · 绿色 被引 0 · OpenAlex

分布匹配蒸馏(DMD)通过分别估计的目标分数与学生分数之差训练少步学生模型,因而必须维护一个拟合学生动态分布的辅助扩散模型,带来额外的显存与计算开销。本文提出 DMAD——分布匹配对抗蒸馏,将分布匹配重构为分类任务,直接学习所需的 log-density 比。共享主干上的两个判别头分别区分真实数据与教师样本是否来自学生模型,并对其 logits 施加线性损失以训练……Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the

Co-Evolving Robot Orchestrators and Policies through Deployment
通过部署协同进化机器人编排器与策略
arXiv:2610.09228 多模态 应用落地

在大规模数据集上训练的视觉-语言-动作(VLA)策略在其训练域内表现良好,但仍难以泛化到真实部署中机器人遭遇的各种场景。Agentic 机器人系统通过视觉-语言模型(VLM)编排器来弥补策略的不足,该编排器学习何时调用策略、如何下达指令、以及何时改用脚本化技能。然而,由于整个系统围绕一个语言可操控性有限的冻结策略构建,编排器只能规避策略的失败却无法真正克服它们。策略成为瓶颈……Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when to call the policy, how to instruct it, and when to use scripted skills instead. However, because the harness is built around a frozen policy that has limited language steerability, the orchestrator can avoid the policy's failures but never overcome them. The policy becomes the bo

SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
SheetSage2:基于合成监督的连贯主旋律谱转录
arXiv:2610.05336 多模态 方法 OA · 绿色 被引 0 · OpenAlex

将音乐转录为人类可读的乐谱需要对节奏、和声、旋律与曲式的整体理解。两大障碍限制了这一目标的实现:标注录音稀缺,以及精确的局部预测仍可能产生不一致的音乐序列。本文提出 SheetSage2,一个统一的音乐转录框架,结合合成数据、任务特定的结构化解码与自回归蒸馏。自动标注的 MIDI 渲染为音频后,可为各类音乐理解任务提供可扩展的监督。任务特定的结构化解码器整合互补的音乐……Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary mu

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis
NAMVIS:下一尺度自回归多视图图像合成
arXiv:2610.04722 多模态 方法 OA · 绿色 被引 0 · OpenAlex

稀疏视图新视角合成是三维内容创作中的核心问题,但基于扩散的方法受迭代去噪限制,多视图生成在推理时开销高昂。本文提出 NAMVIS,一种无扩散框架,将多视图图像合成重新表述为几何条件下的下一尺度自回归过程。NAMVIS 不通过反复去噪生成目标视图,而是通过少量由粗到细的尺度步预测离散视觉 token,并在同一尺度内以及跨目标视图间并行采样所有 token。为锚定此过程……Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor thi

FastOPD: On-Policy Distillation for Lightweight VLA Deployment
FastOPD:面向轻量 VLA 部署的在线蒸馏
arXiv:2610.02832 多模态 应用落地 OA · 绿色 被引 0 · OpenAlex

视觉-语言-动作(VLA)基础模型规模迅速扩大以提升操作性能与泛化能力,但这种规模化带来了高昂的计算成本,使真实世界部署日益困难。现有方法通常通过设计更小的架构或减少基于流(flow-based)策略中的迭代去噪步数来缓解该问题。本文提出 FastOPD,一个从基础到轻量的 VLA 框架,通过高效的在线蒸馏实现大规模 VLA 的实际部署。具体而言,FastOPD 适配流映射(flow map)……Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map