研究库 论文知识库
Papers · organized/paper_cards

论文

329 张论文卡片 · 多模态 · 方法

开放获取 全部 绿色 · 1640
UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing
UltraTex:释放 2K 多视角扩散模型用于 3D 纹理生成
arXiv:2609.23169 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UltraTex,一种用于高分辨率多视图扩散 3D 纹理生成的高效端到端框架,并引入 Background Token Dropping(在 DiT 主干前移除背景 token)与 Block-Sparse Attention(降低前景序列上的注意力计算)。This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision
TAPe+ML:面向多任务计算机视觉的紧凑结构化表示
arXiv:2609.20869 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,将部分建模负担从网络参数转移至结构化输入表示,可在更低的数据、内存与算力需求下支持紧凑的多任务视觉系统。The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.

Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models
用于量子增强扩散语言模型的 Circuit Hypernetwork
arXiv:2609.24657 多模态 方法 被引 0 · S2

提出 HyperQ,在冻结的掩码扩散语言模型上添加 token 条件化的量子残差分支,支持 token 条件化电路发射,作为一种可处理的量子增强语言建模架构方法。HyperQ is introduced, which adds token-conditioned quantum residual branches to a frozen masked-diffusion language model, and support token-conditioned circuit emission as a tractable architectural approach to quantum-augmented language modelling.

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation
所有模态皆平等,但视频更平等:弥合联合视频生成中的跨注意力差距
arXiv:2609.27901 多模态 方法 被引 0 · S2

介绍 RecCAR(Reciprocal Cross-modal Attention Regularization),一种 KL 正则化器,利用已确立的视频到模态对应关系作为固定参考,将较弱的模态到视频对应关系向其对齐。RecCAR, standing for Reciprocal Cross-modal Attention Regularization, is introduced, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it.

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
Spatial-Interactor:通过与可观测物理世界的交互学习空间推理
arXiv:2609.23038 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Spatial-Interactor,一个通过交互训练 VLM 建模物理世界状态转换的框架,并将学习过程组织为三级课程:L1 被动世界状态转换、L2 主动自我状态转换、L3 长时交互轨迹。Spatial-Interactor is introduced, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories.

X-Planner: Event-Structured Task Planning for Embodied Intelligence
X-Planner:面向具身智能的事件结构化任务规划
arXiv:2609.25187 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 X-Planner,一个面向具身推理的规划前端,同时解决监督与表征问题,并描述了规划文本质量与下游执行情况。This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.

Knowledge Pull Requests for Continual Document Authoring
面向持续文档撰写的知识 Pull Request
arXiv:2609.26634 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Knowledge Pull Requests 相比从来源重写或从零重新生成,能整合更多信息并更好地保留已有内容,同时每生成一个 token 增加的信息量最多。Knowledge Pull Requests integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated.

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
ViRDM:驯服表征分布匹配以实现少步长因果视频生成
arXiv:2609.28923 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ViRDM,一种无需教师与评论家网络的视频后训练方案,将三网络蒸馏转化为仅生成器的后训练,在降低 GPU 显存与训练时间的同时提升视频质量。ViRDM, a teacher- and critic-free video post-training recipe that turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality, is introduced.

Parts-of-Speech as Emergent Categories in SAE Latent Space
词性作为 SAE 潜空间中的涌现类别
arXiv:2609.29362 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果显示 SAE 以分布式、类别依赖的形式定位形态句法信息,而非通过原子化的语法特征。The results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.

DeltaWAM: Delta World Action Models for Bimanual Manipulation
DeltaWAM:面向双手操作的 Delta 世界动作模型
arXiv:2609.28811 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DeltaWAM,通过 dense-anchor、sparse-delta 与 action 三条流联合预测视觉 delta 与动作,并设计了三种在表征与计算共享上有所不同的架构;同时开发了 Streaming Delta Memory (SDM),使用紧凑的观测 delta 更新缓存的 anchor 上下文,从而减少繁重的 video-expert 处理。This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing.

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
AV-GRPO:用于联合音视频生成的模态锚定解耦扩散强化学习
arXiv:2609.29816 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 AV-GRPO,一种以模态为锚点的在线 diffusion RL 框架,以及 5DAV,一种解耦的、难度可调节的训练数据集;在 LoRA 与全量 fine-tuning 条件下,其在生成质量、语义对齐和跨模态同步性上均优于 LTX-2.3。This work proposes AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset that outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning.

Learning multiple visual domains with residual adapters
使用残差适配器学习多个视觉域
arXiv:1705.08045 多模态 方法 OA · 绿色 被引 1105 · S2

本文提出一种可调深度网络架构,借助适配器残差模块可在线切换至不同视觉域;同时引入 Visual Decathlon Challenge 基准,用于评估表征同时捕获十个差异显著视觉域的能力,并衡量其跨域均匀识别的能力。This paper develops a tunable deep network architecture that, by means of adapter residual modules, can be steered on the fly to diverse visual domains and introduces the Visual Decathlon Challenge, a benchmark that evaluates the ability of representations to capture simultaneously ten very differentVisual domains and measures their ability to recognize well uniformly.

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
TrackEverything:通过去重 3D 场景表示实现长时稠密跟踪
arXiv:2609.30222 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TrackEverything 是首个在 40 GB GPU 显存内即可在超过 1000 帧的视频中跟踪所有可见点的 3D tracker;本文还提出 3D WAFT,用场景点云内的高效特征采样取代了显存开销巨大的 4D correlation volumes。TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, and 3D WAFT is proposed, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud.

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
FoMo:生成轨迹中的分叉时刻作为感知距离
arXiv:2609.25716 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一条全自动的数据生成流程,可在无需任何人工标注的情况下生成图像对之间的逐点感知距离标签,并证明 diffusion trajectory 与人类视觉系统高度一致。This paper proposes a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation, and demonstrates that the diffusion trajectory aligns well with the human visual system.

Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning
[标题中文] 利用预训练扩散模型与多模态条件增强摄影测量数字表面模型
arXiv:2609.31199 多模态 方法 被引 0 · S2

本文提出一种改进的 Stable Diffusion 3 架构,采用裁剪后的文本流与逐 patch 的归一化策略,使其能在 LiDAR 数据上稳定训练,并实现从自然图像到高程图的迁移;研究表明多模态条件输入可提升高程精度。A modified Stable Diffusion 3 architecture with a pruned text stream and a patch-wise normalization strategy is introduced, enabling stable training on LiDAR data and transfer from natural images to elevation maps, demonstrating that multimodal conditioning improves elevation accuracy.

BoundInk: Boundary-Aware Online Handwriting Generation
[标题中文] BoundInk:面向边界感知的在线手写生成
arXiv:2604.02103 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

BoundInk 是一种以书写者为条件的框架,将字符间的边界视为显式生成单元,既能保留书写者特有的字形外观,又能在完整文本行内改善连接性与字距。BoundInk is introduced, a writer-conditioned framework that treats inter-character boundaries as explicit generation units and preserves writer-specific glyph appearance while improving connectivity and spacing across complete text lines.

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
[标题中文] VLA-Precision:面向视觉-语言-动作模型高效真实世界在线强化学习的非对称协同自举
arXiv:2609.04355 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 VLA-Precision,一个面向真实世界的高效在线 RL 框架,包含 Asymmetric Co-Bootstrapping (ACoB) 算法与 ACoB-Stream 架构;其中 ACoB-Stream 以不变状态解耦与按需流式传输为设计原则,构建了经验-策略的闭环架构,可实现最高 10.9% 的吞吐与计算效率提升。VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture, and ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, deliver up to 10.9% improvements in throughput and computational efficiency.

LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
LightMIS:无逐级解码器的超轻量医学图像分割。
arXiv:2609.28327 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

LightMIS 通过 Scale-Aligned Projection 块将五级编码器的输出对齐至统一分辨率,仅聚合一次,并利用结合 Adaptive Kernel Fusion 与所提 Progressive Receptive Fusion 模块的 Adaptive Fusion Cascade 精炼融合表示。LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade, which combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module.

LastOPD: Taming Collapse in Latent On-Policy Distillation
LastOPD:驯服潜在 On-Policy 蒸馏中的坍缩问题
arXiv:2609.28845 多模态 方法 OA · 绿色 被引 2 · S2

LastOPD 仅在最后一层状态(LM head 共同读取的接口)施加潜在信号,且仅在进入 token 级 OPD 前的 10 步交叉淡入阶段施加,从而保留潜在信号中的有用部分,并在坍缩发生前将 student 交由 token 级监督。LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
InfiniHand:基于第一人称视频的流式世界空间手部运动估计。
arXiv:2609.35743 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

InfiniHand 是一个端到端流式前馈框架,联合估计 MANO 参数、相机轨迹与手部位置,直接从未标定的自我中心视频出发,相比 ViDiHand 在 ARCTIC PA-p 上降低 21%,并显著缓解世界空间漂移。InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, achieving a 21% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift.

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
GeoVerse: 几何潜空间中的世界一致新视角合成
arXiv:2609.35734 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

GeoVerse 是一个在世界一致条件下合成新视角的框架,在预训练 3D 基础模型的几何潜在空间内进行生成,并通过 ControlNet 风格 adapter 注入视频生成模型的外观先验。GeoVerse is a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model via a ControlNet-style adapter.

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
通过交错强化学习在统一模型中学习原生反思
arXiv:2609.35767 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

UMM-Reflection 在统一模型内利用强化学习(RL)完成完整反思轨迹:兄弟轨迹共享同一初始图像,因此组相对优势可比较不同反思策略,而单条轨迹级优势同时更新反思 token 与基于 flow 的修订,避免了逐轮信用分配的组合爆炸。UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment.

FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
FlowTool: 基于流匹配控制图像修图中的工具参数
arXiv:2609.35673 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FlowTool,一个通过 conditional rectified flow 直接建模以输入图像和用户指令为条件的高质量工具参数分布的框架,并显著提升了推理效率。This work introduces FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow, and significantly improves inference efficiency.

Learning from Teacher Continuations at Student States
在学生状态处从教师续写中学习
arXiv:2609.36246 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

OLIVE 通过从持续演化的 student 生成前缀,在离线蒸馏进入平台期后继续提升,同时更好地保持了 student 的通用能力与可塑性。By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student.

Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
我们能信任教师吗?面向自蒸馏的解耦信用方向-幅度
arXiv:2609.34848 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出解耦信用自蒸馏,理论上将信用方向与幅度解耦为两个可靠信号,并据此校准特权教师监督,从而完成策略优化的 step-to-token 信用分配。Decoupled Credit Self-Distillation is introduced, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision, enabling step-to-token credit assignment for policy optimization.

NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
NVAlign:面向连续自回归流匹配文本到语音系统中非言语控制的直接梯度优化
arXiv:2609.31892 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

NVV-SuperBench 与人工听力评测的结果显示,NVAlign 在标签跟随准确率上优于 SFT 与 Flow-GRPO 基线,证明直接对奖励梯度进行优化可提升连续自回归流匹配 TTS 中的非语言控制能力。Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines, demonstrating that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS.

Multimodal 补充候选
arXiv:2606.13578 多模态 方法 OA · 绿色 被引 4 · S2

构建了 RoboGenesis,一个基于仿真的工作流与数据引擎,可从原子技能组合配置好的实验工作流,对 rollout 进行验证与过滤,并跨支持的机器人配置导出结构化演示数据。RoboGenesis is built, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles.

Omni-IO Skills: Harnessing Your Agent Omni-Native
[标题中文] Omni-IO Skills:让你的 Agent 原生支持全模态
arXiv:2609.31847 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,harness 层面的能力组合是构建广泛、可演进 Omni 系统的实用路径,且无需改变宿主 agent 的推理核心。The results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.

Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
[标题中文] Omni-Decision:面向全模态 Agent 的证据账本式规划
arXiv:2607.11433 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Omni-Decision,一个基于证据账本规划的 omni-modal agent:用显式的证据账本替代不断膨胀的对话历史,记录仍缺失的证据、已确认的内容以及记录间的冲突。Omni-Decision is presented, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict.

Improved Distributional Diffusion Models
[标题中文] 改进的分布式扩散模型
arXiv:2609.37147 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种随机少步生成器,其 FID 在采样预算从 4 增加到 50 NFE 时不会下降,且相同配方可迁移到 text-to-image 生成。A stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation.

AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
AutoRef:面向 Agentic 多参考图生成的 Harness 优化
arXiv:2609.35530 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 AutoRef,在保持两侧模型冻结的前提下自动优化 harness:一个 coding Agent 迭代重写 harness 代码,并发现 AutoRef-Harness,可改进开源权重模型 FLUX。This work proposes AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code, and discovers AutoRef-Harness, which improves the open-weight FLUX.

Reasoning with Image Generation
结合图像生成的推理
arXiv:2609.16409 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在包括多视角空间推理和碰撞预测在内的六项多样化视觉推理任务中,ReImaGin 始终优于纯文本推理和专用 vision-tool 基线,提升幅度高达 25%,证明了灵活、可生成的视觉推理的优势。Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
UniEvo-VL:一种用于多模态模型自我改进的 On-policy 自蒸馏训练方案
arXiv:2609.38721 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 UniEvo-VL,一种面向多模态模型的自演化框架,可在 test-time compute 阶段从这种建设性的自纠错反馈中学习,在无外部监督或指导的情况下提升用户使用多模态模型的体验。UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute, is introduced to enhance the user experience when using multimodal models without external supervision or guidance.

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
关键不在图像本身:无关上下文会扰动 VLM 评判且不提供有效信息
arXiv:2609.37863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MIST(Misleading-Image Stress Test):200 个英文句子,每句围绕一个可作比喻或字面理解的短语,配以对齐图像(描绘其读法)、误导图像(描绘相反读法)或无图像三种条件。MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.

Generative Moment Matching Networks
生成矩匹配网络
arXiv:1502.02761 多模态 方法 OA · 绿色 被引 954 · S2

本文提出一种方法,通过多层感知机的一次前馈传播生成独立样本(与近期提出的 GAN 类似),并使用 MMD 学习生成可被解码为样本的 codes。This work forms a method that generates an independent sample via a single feedforward pass through a multilayer perceptron, as in the recently proposed generative adversarial networks, using MMD to learn to generate codes that can then be decoded to produce samples.

Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS:从自回归解码到掩码预测的高效免转录语音克隆
arXiv:2609.38658 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Tacit-TTS 是一个从 IndexTTS2 蒸馏而来的高效无需转录的零样本语音克隆系统,用掩码非自回归生成替换自回归的文本到语义解码,引入无需训练的声学长度估计,并通过 ReFlow 蒸馏加速流匹配渲染器。Tacit-TTS is presented, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation.