研究库 论文知识库
Papers · organized/paper_cards

论文

31 张论文卡片 · 多模态 · 应用落地 · OA 绿色

开放获取 全部 绿色 · 1640
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
Embodied-Navigator:用于高效导航的指向、思考、记忆与对齐
arXiv:2608.17512 多模态 应用落地 OA · 绿色 被引 1 · S2

该工作提出 TAMP-Nav,一个用于高效具身导航的统一框架,可在关键节点动态触发 Chain-of-Thought 并仅在关键节点保留高保真记忆,将冗余轨迹压缩为轻量级时空指示器,从而保留关键历史信息并增强时空感知。This work proposes TAMP-Nav, a unified framework for efficient embodied navigation that dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception.

PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage
PinSieve:生产级选择性 VLM 服务与企业内容质量分诊的可治理记忆飞轮
arXiv:2608.24040 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 PinSieve——大规模内容质量流水线中的生产级案例:一个选择性 vision-language-model Serving Agent,仅处理轻量上游模型无法覆盖的 grey-zone 切片,在线暴露标量路由评分,并保留受控的人工升级通道。This work presents PinSieve, a production case study in a large-scale content-quality pipeline, a selective vision-language-model Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation.

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
缺失的时间纽带:面向剧本驱动音视频生成的时间上下文路由
arXiv:2609.02367 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Temporal Context Routing,将脚本时序映射到视频与音频生成的共享时间轴上,并将每个 prompt 的引导路由到两种模态中的对应位置,同时保持与 baseline 相当的视觉质量与音视频同步性。Temporal Context Routing is introduced, which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities, while maintaining visual quality and audio-visual synchronization comparable to those of the baselines.

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
ShallowStream:先浅层索引再深层回答的流式视频理解
arXiv:2609.02780 多模态 应用落地 OA · 绿色 被引 2 · S2

ShallowStream 是一个利用 MLLM 浅层同时进行帧编码与检索索引构建的新框架,性能与当前最强流式方法相当,同时将单帧 prefill 延迟与 10 秒端到端延迟分别降低至多 52.1 倍和 11.9 倍。ShallowStream is a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building and achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively.

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
EmbodiedSkills:用于编排、训练和部署 VLA Agent 的统一框架
arXiv:2609.01281 多模态 应用落地 OA · 绿色 被引 2 · S2

EmbodiedSkills 是一个统一框架,将每项技能决策视为执行提案:运行时先检查前置条件再执行,执行后再验证结果,为将底层 VLA 策略转化为闭环具身系统提供可训练且可检视的 Agent 层。EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward and provides a trainable and inspectable agent layer for turning low-level VLA policies into closed-loop embodied systems.

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Marigold V2:重访用于单目深度估计的扩散 Transformer
arXiv:2609.08084 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

Marigold V2 在应用于其他稠密回归任务(如表面法向量估计与本征图像分解)时取得 SOTA 结果。Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
基于合成语音构建并评估固定音色的泰语 TTS
arXiv:2609.03502 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

最终得到的 82M 参数模型 Wayu-Paxa-TTS-Edge 实现了无需参考音频的设备端泰语 TTS,并在三个系统中取得了最低的停顿位置错误率与词内停顿率。The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
FRAUDSkill:面向音频反诈骗检测的结构化冻结权重 Skill 优化
arXiv:2609.18766 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FRAUDSkill——一种结构化的 frozen-weight 适配框架:底层音频-语言模型保持不变,转而优化外部的 skill program、路由策略和决策规则,并将结构化输出控制与验证引导的多路径推理相结合,以确保符合协议规范的预测。FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions.

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads
MLLMs 在 Synergy Heads 中信息分布漂移时产生幻觉
arXiv:2609.09206 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

HEAL 将动态信息校准因子注入协同注意力头的 value 向量中,主动调节视觉-语言依赖,引导输出分布趋向事实证据,提供了一条简单且可解释的增强模型可信度的路径。HEAL injects dynamic information calibration factors into the value vectors of synergy heads and actively regulates visual-language dependencies, steering the output distribution towards factual evidence, offering a simple and interpretable pathway to enhance model trustworthiness.

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
脚本感知的 Mixture-of-Experts 统一多语言场景文本识别
arXiv:2609.24058 多模态 应用落地 OA · 绿色 被引 1 · S2

本文构建了 TextMuSS-10M,一个涵盖 10 种文字、229 种语言的大规模合成场景文本数据集,并提出了 ScriptMoE,一种具备文字感知能力的 Mixture-of-Experts (MoE) 架构。该架构在精度上达到最高,且比 per-language experts 更简单、比 VLM 更轻量,同时精度优于两者。This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.

WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation
[标题中文] WISE-ATTA:预算受限的主动测试时自适应中的标签请求时机
arXiv:2609.37687 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Budgeted ATTA:测试批次中仅有一小部分可获得标签,且监督时机是 active test-time adaptation 中一个关键但尚未充分探索的方面。Budgeted ATTA is introduced in which labels are available for only a fraction of test batches, and the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation.

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing
在机器人执行误差下驯服 VLA:自补偿与压力测试
arXiv:2609.37334 多模态 应用落地 OA · 绿色 被引 1 · S2

论文提出 Self-compensating VLA,一种部署阶段的自适应方法,使 VLA 策略在生成指令时能够预补偿机器人的执行误差,并在平均任务成功率上高于基线策略以及在训练阶段增强鲁棒性的方法。Self-compensating VLA is proposed, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands, and achieves higher average task success than both the base policies and methods that build in robustness during training.

FastOPD: On-Policy Distillation for Lightweight VLA Deployment
FastOPD:面向轻量 VLA 部署的在线蒸馏
arXiv:2610.02832 多模态 应用落地 OA · 绿色 被引 0 · OpenAlex

视觉-语言-动作(VLA)基础模型规模迅速扩大以提升操作性能与泛化能力,但这种规模化带来了高昂的计算成本,使真实世界部署日益困难。现有方法通常通过设计更小的架构或减少基于流(flow-based)策略中的迭代去噪步数来缓解该问题。本文提出 FastOPD,一个从基础到轻量的 VLA 框架,通过高效的在线蒸馏实现大规模 VLA 的实际部署。具体而言,FastOPD 适配流映射(flow map)……Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map

From Foundation to Application: Improving VLA Models in Practice
从基础到应用:实践中改进 VLA 模型
arXiv:2607.06403 多模态 应用落地 OA · 绿色 被引 28 · S2

得益于涵盖全身自由度的扩展预训练数据,LingBot-VLA-2.0 在两个机器人平台上展现出强大的跨具身长时程移动操作能力。Benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
MultAttnAttrib:长文档问答中的免训练多模态归因
arXiv:2607.01420 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 MultAttnAttrib,一种免训练的归因生成方法,利用模型的预填充过程、选定的注意力头以及校准阈值在文档中定位源证据,且在多种归因生成方法上一致地表现更优。This work introduces MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a document, and consistently outperforms a variety of attribution-generation methods.

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
京东 Oxygen AI 商品中心(Oxygen AIIC)V1:以 LLM/VLM 为核心的工业级商品理解、管理与应用解决方案
arXiv:2606.28070 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

京东 Oxygen AI 商品中心(Oxygen AIIC):基于 LLM/VLM 的工业级商品知识生产与服务平台,已在大规模场景下取得可量化的收益。The JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service, has delivered measurable gains at scale.

Vesta: A Generalist Embodied Reasoning Model
Vesta:通用具身推理模型
arXiv:2606.20905 多模态 应用落地 OA · 绿色 被引 6 · S2

提出 Vesta,一个统一的具身通用模型,将定位、空间推理、导航和长程规划能力整合到单个基础模型中,并证明通用模型能够达到或超越专家模型。Vesta is presented, a unified embodied generalist that consolidates localization, spatial reasoning, navigation, navigation, and long-horizon planning capabilities into a single foundation model and demonstrates that a generalist model can match or exceed specialists.

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
基于物体中心的残差强化学习用于 VLA 零样本仿真到现实迁移的增强
arXiv:2606.18953 多模态 应用落地 OA · 绿色 被引 1 · S2

提出一种基于物体中心的残差强化学习框架,利用物体位姿精化 VLA 动作,使观测空间紧凑,在仿真与现实之间能够一致迁移。An object-centric residual RL framework is proposed that refines VLA actions using object poses, enabling a compact observation space that transfers consistently between simulation and reality.

A Quantized Native Runtime for On-Device Semantic Audio Generation
一种用于端侧语义音频生成的量化原生运行时
arXiv:2607.08526 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

该紧凑且量化原生、带内置控制的运行时为物联网音频场景下的端侧语义音频提供了实用基础;通过对转向接口的案例分析,可生成在部分属性上具有真实但有界控制的、承载口味联想的音乐。A compact, quantized runtime with built-in control a practical basis for on-device semantic audio in Internet-of-Sounds settings and a case study of the steering interface generates music carrying taste associations with genuine but bounded control for a subset of attributes.

Motion4Motion: Motion Transfer Across Subjects at Inference
Motion4Motion:推理阶段的跨主体运动迁移
arXiv:2607.11644 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

Motion4Motion 对视频中角色的 motion flow(而非骨骼)进行建模,使跨物种运动迁移更加容易。Motion4Motionmodels the motion flow of the character in a video instead of skeletons, which makes motion transfer across species easier, which makes motion transfer across species easier.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Xiaomi-Robotics-U0:基于 World Foundation Model 的统一具身合成
arXiv:2607.11643 多模态 应用落地 OA · 绿色 被引 4 · S2

Xiaomi-Robotics-U0 是首个支持跨多种机器人本体的高质量多视角场景生成、并引入结构化、可控的具身迁移以实现细粒度编辑的模型,同时保持多视角一致性与交互动态。Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
VideoChat3:面向高效通用视频理解的完全开放视频 MLLM
arXiv:2607.14935 多模态 应用落地 OA · 绿色 被引 8 · S2

本文提出 VideoChat3,一个完全开放、高效且以视频为中心的通用 MLLM,仅以 4B 参数与更高效率,超越参数量相当或更大的已有开源模型。This work introduces VideoChat3, a fully open, efficient, and generalist video-centric MLLM, which surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.

Robostral Navigate
Robostral Navigate
arXiv:2607.20785 多模态 应用落地 OA · 绿色 被引 2 · S2

提出 Robostral Navigate,一个围绕该可扩展性目标构建的 8B 视觉语言模型,仅消费单目 RGB 图像流——这是机器人平台中最普及的传感器——通过在当前相机画面中指向下一目标位置来预测航点。Robostral Navigate, an 8B vision-language model built around this scalability objective, is introduced, which consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view.

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
Closing the Loop:面向自回归生成式渲染的无训练 Revisit 一致性
arXiv:2607.21848 多模态 应用落地 OA · 绿色 被引 0 · OpenAlex

近期条件视频生成模型已展现出将 3D 引擎渲染(如深度图与无纹理几何体)转化为照片级真实视频的潜力,可应用于游戏与沉浸式内容创作。此类应用要求长时程自回归生成,在持续合成新帧的同时维持持久的 3D 世界。自回归生成器以有界 KV cache 逐 chunk 合成视频,因此当相机再次访问已从上下文中驱逐的位置时,模型常会重新生成不一致的外观,尽管该Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Auto-regressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after its context has been evicted, the model often regenerates inconsistent appearance, even though the

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
UltraViT:面向大视觉-语言模型的端侧延迟优化视觉编码器
arXiv:2607.23373 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

大量实验表明,结合端上延迟感知设计与定制化训练策略,建立了高效 LVLM 编码的新 SOTA,在端上以近 1.7 倍速度运行的同时显著优于现有以编码器为中心的基线。Extensive experiments demonstrate that the on-device latency-informed design combined with the tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.

Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection
UAV 高光谱 PFM-1 地雷检测中的人在回路签名引导
arXiv:2607.25310 多模态 应用落地 OA · 绿色 被引 1 · S2

本文研究无人机(UAV)可见光-近红外(VNIR)高光谱图像中 PFM-1 地雷的检测,使用光谱角制图(SAM)、匹配滤波器(MF)、自适应相干估计器(ACE)和约束能量最小化(CEM)。This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using spectral angle mapper (SAM), matched filter (MF), adaptive coherence estimator (ACE), and constrained energy minimization (CEM).

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
标题 -> 标题中文:响亮还是沉默?面向多模态临床 AI 的可复用逐模态失效分析框架
arXiv:2608.01462 多模态 应用落地 OA · 绿色 被引 1 · S2

一种模型无关的模态失败框架,仅依赖部署可观察信号,返回逐样本失败分类、逐模态互补矩阵(将错误归因到模态)以及响亮 vs 静默 dropout 画像(区分可监控失败与远离决策边界未被标记的失败)A model-agnostic modality-failure framework that returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals.

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
BridgeVLA++:用于 3D 操作的、数据高效、可泛化且具备记忆增强的 Vision-Language-Action 框架
arXiv:2608.05042 多模态 应用落地 OA · 绿色 被引 4 · S2

在 BridgeVLA 基础上开发 BridgeVLA++,引入统一的时空记忆架构,建模持久化的空间上下文与时间交互历史,使其可在保留 BridgeVLA 数据效率与泛化能力的同时对观测历史进行推理。BridgeVLA++ is developed by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history that can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities.

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
往返一致性:双向扩散模型可预测自身的 rollout 误差
arXiv:2608.00675 多模态 应用落地 OA · 绿色 被引 1 · S2

往返一致性将可逆性转化为生成式模型一种实用的可信信号;双向训练带来负成本,在两个方向上均优于单向专家模型;其中反向还可作为快速的逆问题求解器。Round-trip consistency turns reversibility into a practical trust signal for generative models, and Bidirectional training comes at negative cost, beating direction specialists in both directions, and the backward direction doubles as a fast inverse solver.

Co-occurrence Feature Learning for Skeleton based Action Recognition using Regularized Deep LSTM Networks
使用正则化深度 LSTM 网络进行基于骨骼动作识别的共现特征学习
arXiv:1603.07772 多模态 应用落地 OA · 绿色 被引 930 · S2

本文在每个时间步以骨骼作为输入,引入一种新的正则化方案来学习骨骼关节的共现特征,并提出一种同时作用于 LSTM 神经元门、单元和输出响应的新型 dropout 算法。This work takes the skeleton as the input at each time slot and introduces a novel regularization scheme to learn the co-occurrence features of skeleton joints, and proposes a new dropout algorithm which simultaneously operates on the gates, cells, and output responses of the LSTM neurons.

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
LiveAnimate:实时稳定的长视频流式人体动画
arXiv:2608.11745 多模态 应用落地 OA · 绿色 被引 1 · S2

本文提出 LiveAnimate,据作者所知是首个将实时流式生成与十亿参数规模下的稳定长视频生成相结合的系统,基于 140 亿参数的视频 Diffusion Transformer(DiT)。This work presents LiveAnimate, to their knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT).