提出 UltraTex,一种用于高分辨率多视图扩散 3D 纹理生成的高效端到端框架,并引入 Background Token Dropping(在 DiT 主干前移除背景 token)与 Block-Sparse Attention(降低前景序列上的注意力计算)。This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.
论文
329 张论文卡片 · 多模态 · 方法
结果表明,将部分建模负担从网络参数转移至结构化输入表示,可在更低的数据、内存与算力需求下支持紧凑的多任务视觉系统。The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.
提出 HyperQ,在冻结的掩码扩散语言模型上添加 token 条件化的量子残差分支,支持 token 条件化电路发射,作为一种可处理的量子增强语言建模架构方法。HyperQ is introduced, which adds token-conditioned quantum residual branches to a frozen masked-diffusion language model, and support token-conditioned circuit emission as a tractable architectural approach to quantum-augmented language modelling.
介绍 RecCAR(Reciprocal Cross-modal Attention Regularization),一种 KL 正则化器,利用已确立的视频到模态对应关系作为固定参考,将较弱的模态到视频对应关系向其对齐。RecCAR, standing for Reciprocal Cross-modal Attention Regularization, is introduced, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it.
介绍 Spatial-Interactor,一个通过交互训练 VLM 建模物理世界状态转换的框架,并将学习过程组织为三级课程:L1 被动世界状态转换、L2 主动自我状态转换、L3 长时交互轨迹。Spatial-Interactor is introduced, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories.
本文提出 X-Planner,一个面向具身推理的规划前端,同时解决监督与表征问题,并描述了规划文本质量与下游执行情况。This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.
Knowledge Pull Requests 相比从来源重写或从零重新生成,能整合更多信息并更好地保留已有内容,同时每生成一个 token 增加的信息量最多。Knowledge Pull Requests integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated.
本文提出 ViRDM,一种无需教师与评论家网络的视频后训练方案,将三网络蒸馏转化为仅生成器的后训练,在降低 GPU 显存与训练时间的同时提升视频质量。ViRDM, a teacher- and critic-free video post-training recipe that turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality, is introduced.
结果显示 SAE 以分布式、类别依赖的形式定位形态句法信息,而非通过原子化的语法特征。The results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
本文提出 DeltaWAM,通过 dense-anchor、sparse-delta 与 action 三条流联合预测视觉 delta 与动作,并设计了三种在表征与计算共享上有所不同的架构;同时开发了 Streaming Delta Memory (SDM),使用紧凑的观测 delta 更新缓存的 anchor 上下文,从而减少繁重的 video-expert 处理。This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing.
本文提出 AV-GRPO,一种以模态为锚点的在线 diffusion RL 框架,以及 5DAV,一种解耦的、难度可调节的训练数据集;在 LoRA 与全量 fine-tuning 条件下,其在生成质量、语义对齐和跨模态同步性上均优于 LTX-2.3。This work proposes AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset that outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning.
本文提出一种可调深度网络架构,借助适配器残差模块可在线切换至不同视觉域;同时引入 Visual Decathlon Challenge 基准,用于评估表征同时捕获十个差异显著视觉域的能力,并衡量其跨域均匀识别的能力。This paper develops a tunable deep network architecture that, by means of adapter residual modules, can be steered on the fly to diverse visual domains and introduces the Visual Decathlon Challenge, a benchmark that evaluates the ability of representations to capture simultaneously ten very differentVisual domains and measures their ability to recognize well uniformly.
TrackEverything 是首个在 40 GB GPU 显存内即可在超过 1000 帧的视频中跟踪所有可见点的 3D tracker;本文还提出 3D WAFT,用场景点云内的高效特征采样取代了显存开销巨大的 4D correlation volumes。TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, and 3D WAFT is proposed, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud.
本文提出一条全自动的数据生成流程,可在无需任何人工标注的情况下生成图像对之间的逐点感知距离标签,并证明 diffusion trajectory 与人类视觉系统高度一致。This paper proposes a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation, and demonstrates that the diffusion trajectory aligns well with the human visual system.
本文提出一种改进的 Stable Diffusion 3 架构,采用裁剪后的文本流与逐 patch 的归一化策略,使其能在 LiDAR 数据上稳定训练,并实现从自然图像到高程图的迁移;研究表明多模态条件输入可提升高程精度。A modified Stable Diffusion 3 architecture with a pruned text stream and a patch-wise normalization strategy is introduced, enabling stable training on LiDAR data and transfer from natural images to elevation maps, demonstrating that multimodal conditioning improves elevation accuracy.
BoundInk 是一种以书写者为条件的框架,将字符间的边界视为显式生成单元,既能保留书写者特有的字形外观,又能在完整文本行内改善连接性与字距。BoundInk is introduced, a writer-conditioned framework that treats inter-character boundaries as explicit generation units and preserves writer-specific glyph appearance while improving connectivity and spacing across complete text lines.
本文提出 VLA-Precision,一个面向真实世界的高效在线 RL 框架,包含 Asymmetric Co-Bootstrapping (ACoB) 算法与 ACoB-Stream 架构;其中 ACoB-Stream 以不变状态解耦与按需流式传输为设计原则,构建了经验-策略的闭环架构,可实现最高 10.9% 的吞吐与计算效率提升。VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture, and ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, deliver up to 10.9% improvements in throughput and computational efficiency.
LightMIS 通过 Scale-Aligned Projection 块将五级编码器的输出对齐至统一分辨率,仅聚合一次,并利用结合 Adaptive Kernel Fusion 与所提 Progressive Receptive Fusion 模块的 Adaptive Fusion Cascade 精炼融合表示。LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade, which combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module.
LastOPD 仅在最后一层状态(LM head 共同读取的接口)施加潜在信号,且仅在进入 token 级 OPD 前的 10 步交叉淡入阶段施加,从而保留潜在信号中的有用部分,并在坍缩发生前将 student 交由 token 级监督。LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.
InfiniHand 是一个端到端流式前馈框架,联合估计 MANO 参数、相机轨迹与手部位置,直接从未标定的自我中心视频出发,相比 ViDiHand 在 ARCTIC PA-p 上降低 21%,并显著缓解世界空间漂移。InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, achieving a 21% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift.
GeoVerse 是一个在世界一致条件下合成新视角的框架,在预训练 3D 基础模型的几何潜在空间内进行生成,并通过 ControlNet 风格 adapter 注入视频生成模型的外观先验。GeoVerse is a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model via a ControlNet-style adapter.
UMM-Reflection 在统一模型内利用强化学习(RL)完成完整反思轨迹:兄弟轨迹共享同一初始图像,因此组相对优势可比较不同反思策略,而单条轨迹级优势同时更新反思 token 与基于 flow 的修订,避免了逐轮信用分配的组合爆炸。UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment.
本文提出 FlowTool,一个通过 conditional rectified flow 直接建模以输入图像和用户指令为条件的高质量工具参数分布的框架,并显著提升了推理效率。This work introduces FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow, and significantly improves inference efficiency.
OLIVE 通过从持续演化的 student 生成前缀,在离线蒸馏进入平台期后继续提升,同时更好地保持了 student 的通用能力与可塑性。By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student.
提出解耦信用自蒸馏,理论上将信用方向与幅度解耦为两个可靠信号,并据此校准特权教师监督,从而完成策略优化的 step-to-token 信用分配。Decoupled Credit Self-Distillation is introduced, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision, enabling step-to-token credit assignment for policy optimization.
NVV-SuperBench 与人工听力评测的结果显示,NVAlign 在标签跟随准确率上优于 SFT 与 Flow-GRPO 基线,证明直接对奖励梯度进行优化可提升连续自回归流匹配 TTS 中的非语言控制能力。Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines, demonstrating that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS.
构建了 RoboGenesis,一个基于仿真的工作流与数据引擎,可从原子技能组合配置好的实验工作流,对 rollout 进行验证与过滤,并跨支持的机器人配置导出结构化演示数据。RoboGenesis is built, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles.
结果表明,harness 层面的能力组合是构建广泛、可演进 Omni 系统的实用路径,且无需改变宿主 agent 的推理核心。The results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.
提出 Omni-Decision,一个基于证据账本规划的 omni-modal agent:用显式的证据账本替代不断膨胀的对话历史,记录仍缺失的证据、已确认的内容以及记录间的冲突。Omni-Decision is presented, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict.
一种随机少步生成器,其 FID 在采样预算从 4 增加到 50 NFE 时不会下降,且相同配方可迁移到 text-to-image 生成。A stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation.
本文提出 AutoRef,在保持两侧模型冻结的前提下自动优化 harness:一个 coding Agent 迭代重写 harness 代码,并发现 AutoRef-Harness,可改进开源权重模型 FLUX。This work proposes AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code, and discovers AutoRef-Harness, which improves the open-weight FLUX.
在包括多视角空间推理和碰撞预测在内的六项多样化视觉推理任务中,ReImaGin 始终优于纯文本推理和专用 vision-tool 基线,提升幅度高达 25%,证明了灵活、可生成的视觉推理的优势。Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.
本文提出 UniEvo-VL,一种面向多模态模型的自演化框架,可在 test-time compute 阶段从这种建设性的自纠错反馈中学习,在无外部监督或指导的情况下提升用户使用多模态模型的体验。UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute, is introduced to enhance the user experience when using multimodal models without external supervision or guidance.
本文提出 MIST(Misleading-Image Stress Test):200 个英文句子,每句围绕一个可作比喻或字面理解的短语,配以对齐图像(描绘其读法)、误导图像(描绘相反读法)或无图像三种条件。MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.
本文提出一种方法,通过多层感知机的一次前馈传播生成独立样本(与近期提出的 GAN 类似),并使用 MMD 学习生成可被解码为样本的 codes。This work forms a method that generates an independent sample via a single feedforward pass through a multilayer perceptron, as in the recently proposed generative adversarial networks, using MMD to learn to generate codes that can then be decoded to produce samples.
Tacit-TTS 是一个从 IndexTTS2 蒸馏而来的高效无需转录的零样本语音克隆系统,用掩码非自回归生成替换自回归的文本到语义解码,引入无需训练的声学长度估计,并通过 ReFlow 蒸馏加速流匹配渲染器。Tacit-TTS is presented, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation.