研究库 论文知识库
Papers · organized/paper_cards

论文

443 张论文卡片 · 多模态

开放获取 全部 绿色 · 1640
How Far Can Synthetic Data Take Thai OCR?
合成数据能将泰语 OCR 带到多远?
arXiv:2609.03595 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现,仅使用合成数据的训练即可具备竞争力;将 0.9B 参数的 PaddleOCR-VL-1.6 适配为 Wayu-Paxa-OCR-Zero,一个无需真实泰语文档页 OCR 标签即可适配的泰语 OCR 模型,表明仅合成数据训练即可具备竞争力。It is found that synthetic-only training can be competitive, and the 0.9B-parameter PaddleOCR-VL-1.6 is adapted into Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages, showing that synthetic-only training can be competitive.

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
arXiv:2609.12036 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Pelican-Sim 在轨迹、场景、物体、具身和视角变化上的定性泛化能力,凸显其作为通用世界模型模拟器的潜力。Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights Pelican-Sim's potential as a general-purpose world model simulator.

Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision
Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision
arXiv:2609.07099 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提交于 ECCV 2026 Wearable AI Challenge 的 EgoProactive 赛道,在大模型组排名第一、≤2B 组排名第二,表明对当前任务而言视觉定位比标注量更为重要。This submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge is presented, which ranked first in the large-model division and second in the<=2B division, suggesting that visual grounding is more important than annotation volume for this task.

Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model
Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model
arXiv:2609.07154 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该系统仅用一个 2B 视觉-语言模型,通过一次贪心前向传播即可回答关于十分钟第一人称视频的多选题,并仅用大型 Agentic pipeline 1.1% 的参数即达到其 89% 的准确率。The system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass, and reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters.

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
arXiv:2609.10895 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

对突发物理危险的反应既是具身智能的重要检验,也是将多模态大语言模型 (MLLMs) 部署为家庭机器人决策核心的硬性要求;ReactHuman 是首个面向类人反应式决策的物理驱动基准。Reacting to sudden physical hazards is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision core of household robots, and ReactHuman is the first physics-grounded benchmark for human-like reactive decision-making.

MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control
MobileVLA-R1 2.0:面向移动机器人控制的 RL 增强推理
arXiv:2609.06251 多模态 方法 OA · 绿色 被引 1 · S2

提出 MobileVLA-R1 2.0,一个 RL 增强的 VLA 框架,显式地将结构化具身推理与可执行的移动机器人控制耦合,并引入推理条件化的动作解码器,将多模态推理表征映射到任务级动作目标,再由机器人控制器翻译为具身特定的指令。This work proposes MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control, and introduces a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers.

Realtime-Venus: A full-duplex interaction system with asynchronous delegation
Realtime-Venus:具备异步委托能力的全双工交互系统
arXiv:2609.13814 多模态 方法 OA · 绿色 被引 2 · S2

Realtime-Venus 是一个主动全双工交互系统,由两个独立训练的 9B 模型组成:Realtime-Venus-Omni 用于音视频交互,Realtime-Venus-Audio 用于语音交互;在 MMAU、Llama Questions 与 Speech CMMLU 上均领先于对比模型Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, which leads the compared models on MMAU, Llama Questions, and Speech CMMLU.

LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
LynnReal-Omni:面向Agent视觉工作流的原生多模态视频生成
arXiv:2609.15863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

视频扩散模型具有随机性且难以控制:精确内容往往需要反复采样且无法保证成功,长时场景在外观、交互和时间一致性上会发生漂移。Agent式视觉创作可提供显式参考、可编辑的3D场景或可执行的游戏状态以实现稳定控制,但本身并不能保证高对象或角色保真度。二者结合可实现稳定且高质量的生成。为实现该结合,我们提出 LynnReal-Omni,一个基于32B共享多模态Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or executable game states for stable control, but does not by itself guarantee high object or character fidelity. Combining the two can enable stable, high-quality generation. To realize this combination, we present LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimo

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
ModaLens:报告条件医学 VLM 中图像敏感性的度量
arXiv:2609.15635 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ModaLens 是一种配对图像交换审计,用于衡量报告可用性如何改变图像敏感性,并限制对视觉正确性的结论;该方向在另外两个模型系列中得到复现。ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity and limits conclusions about visual correctness; the direction replicates in two further model lineages.

Building a Production Greek-English Speech Recognizer
构建生产级希腊语-英语语音识别器
arXiv:2609.13498 多模态 观点 被引 0 · S2

我们报告了一项历时数月的工程项目,用于构建 Sophea,一个生产级希腊语-英语双语自动语音识别系统。我们依据九个生产门控评估该系统,涵盖希腊语和英语词错误率、语言识别及非语音音频的幻觉问题。经过二十三次训练迭代和两种模型架构,没有任何训练数据组合能同时通过全部九个门控。满足希腊语噪声环境目标需要约1,500步密集领域暴露,而保持英语语言识别仅能容忍约2%We report a multi-month engineering program to build Sophea, a production bilingual Greek-English automatic speech recognition system. We evaluate the system against nine production gates covering Greek and English word error rate, language identification, and hallucinations on non-speech audio. Across twenty-three training iterations and two model architectures, no training-data composition passed all nine gates simultaneously. Meeting the Greek noisy-environment target required about 1,500 steps of dense domain exposure, while preserving English language identification tolerated only about 2

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Dynin-Robotics:全模态统一扩散视觉-语言-动作模型
arXiv:2609.13053 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Dynin-Robotics 在 LIBERO 和零样本 LIBERO-Plus 上取得了具有竞争力的性能,并在 Franka Research 3 机器人的四种操作条件下达到了 78.4% 的平均成功率。Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
PhysStream:基于结构化场景记忆与细粒度运动控制的流式物理约束视频生成
arXiv:2609.17521 多模态 观点 OA · 绿色 被引 1 · S2

PhysStream 是一个用于物理驱动图像到视频合成的自回归模型,通过引入结构化场景记忆并支持基于稀疏速度增量信号的细粒度运动控制(编码物理量),使模型能够学习底层动力学。PhysStream is an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics.

StepAudio 3 Realtime Technical Report
StepAudio 3 Realtime 技术报告
arXiv:2609.14005 多模态 方法 OA · 绿色 被引 3 · S2

StepAudio 3 Realtime 是一个围绕持续 listen-converse-think-act 循环组织的音频-语言基础模型,在实时语音输出的同时达到与专用推理模型相当的对话与推理性能,并通过 Think-While-Speaking 机制化解深度思考与延迟之间的矛盾。StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop, achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time, and resolves the tension between deep deliberation and latency via Think-While-Speaking.

StepAudio 3 Music Technical Report
StepAudio 3 Music 技术报告
arXiv:2609.16034 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了 StepAudio 3 Music,一个支持显式音乐规划与开放域文本控制生成的大规模长篇音乐生成模型,在所评估系统中取得最高的 AudioBox Content Enjoyment、Content Usefulness 与 Production Quality 分数,以及最高的 MuQ-MuLan 相似度。This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems.

Register Tokens for Bounded-State Reasoning in Diffusion Language Models
用于扩散语言模型有界状态推理的 Register Tokens
arXiv:2609.16372 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

Register被实现为专用的固定位置 token,其连续的隐藏状态被训练用于跨生成块承载推理进度,对有界代码生成尤其有效,因为正确程序通常跨越多个块。Registers are implemented as dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks, which are especially effective for bounded code generation, where correct programs usually span several chunks.

Convergent Emergence of In-Context Learning Across Modalities
跨模态上下文学习的会聚涌现
arXiv:2609.14011 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个受控的跨模态框架,在多种模态下实例化相同的任务套件以检验 Convergent Emergence Hypothesis,结果显示配对映射的 ICL 在六种模态中出现,超越受控基线,并在其中五种模态上呈现相关的逐任务效应。A controlled cross-modality framework that instantiates the same task suite in a variety of modalities to test the Convergent Emergence Hypothesis and shows that paired-mapping ICL emerges across six modalities, surpasses controlled baselines, and has correlated per-task effects across five of them.

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
ActionPiece: 重新思考自回归视觉-语言-动作模型中的动作 token 化
arXiv:2609.18487 多模态 观点 OA · 绿色 被引 1 · S2

本文提出物理秩一致性 (PRC) 来衡量 tokenization 在重建后保留局部物理距离排序的程度,并提出 ActionPiece,通过对表示学习和量化的联合监督来保留物理动作关系。This work introduces physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction, and presents ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization.

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
Zing-0.5: 迈向具备实时联合动作与文本控制的可玩世界
arXiv:2609.17909 多模态 方法 OA · 绿色 被引 2 · S2

我们提出 Zing-0.5,一个 5B 自回归世界模型,专注于可玩性:用户可以探索生成的世界、影响事件演进,并通过键盘与在线文本的联合控制对反馈做出响应。我们的方法汇聚了三项技术贡献:(1) 统一的动作与文本条件建模,将感知幅度的键盘输入与时序对齐的文本指令、以及联合标注的视频结合,在同一序列中学习导航与事件控制;(2) 面向增量生成的事件尺度监督,使用分段级教师...We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trai

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
PANORAMA:基于掩码提议选择的全景接地描述生成
arXiv:2609.19143 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PanoCaps,一个由全景分割数据集构建的人工标注 benchmark,并提出 PANORAMA,一个将预训练 segmenter 条件化于上下文化短语表示以获得候选 mask、并学习选择与每个短语对应的 mask 的 VLM。This work introduces PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets, and introduces PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase.

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
UFO:面向多模态图像生成全条件对齐的链式评估
arXiv:2609.12397 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UFO,这是首个面向全条件对齐同时评估的统一框架;并发布 UFO-Bench,一个用于整体评估现有定制化模型在文本与视觉条件多样化交互下表现的专用基准。UFO is proposed, the first unified framework for omni-condition alignment simultaneous evaluation, and UFO-Bench is presented, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
FAMOS:基于稀疏观测的前馈 3D 关节建模
arXiv:2609.20817 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FAMOS,一个前馈模型,可从稀疏、无序的部分点云集合预测可动部件分割与关节参数;并引入一个过程式数据生成器,在训练过程中合成自标注资产,以克服现有数据集规模和多样性的局限。FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds, is presented and a procedural data generator that synthesizes self-annotated assets during training is introduced to overcome the limited scale and diversity of existing datasets.

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
MiniMax-H3 能否对物理世界进行推理?一种全模态生成模型的评估
arXiv:2609.18323 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个围绕物理世界推理四个互补维度构建的综合评估框架,并构建了一系列多样化新任务,要求模型跨模态整合互补信息。This work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning, and constructs a diverse set of novel tasks that require models to integrate complementary information across modalities.

Paint-Anything: Unified Any-Color Control for Image Generation and Editing
Paint-Anything:面向图像生成与编辑的统一任意颜色控制
arXiv:2609.20816 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Paint-Anything,通过物体级颜色监督学习一个统一的 hex-prompt 界面以同时支持生成与编辑;并引入 Any Color Benchmark (ACBench),包含 ACBench-T2I 和 ACBench-Edit,以衡量两项任务中的物体级 hex 颜色保真度。Paint-Anything is presented, which learns a shared hex-prompt interface for generation and editing through object-level color supervision, and introduces Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks.

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
FRAUDSkill:面向音频反诈骗检测的结构化冻结权重 Skill 优化
arXiv:2609.18766 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FRAUDSkill——一种结构化的 frozen-weight 适配框架:底层音频-语言模型保持不变,转而优化外部的 skill program、路由策略和决策规则,并将结构化输出控制与验证引导的多路径推理相结合,以确保符合协议规范的预测。FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions.

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads
MLLMs 在 Synergy Heads 中信息分布漂移时产生幻觉
arXiv:2609.09206 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

HEAL 将动态信息校准因子注入协同注意力头的 value 向量中,主动调节视觉-语言依赖,引导输出分布趋向事实证据,提供了一条简单且可解释的增强模型可信度的路径。HEAL injects dynamic information calibration factors into the value vectors of synergy heads and actively regulates visual-language dependencies, steering the output distribution towards factual evidence, offering a simple and interpretable pathway to enhance model trustworthiness.

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
OmniVChat:面向原生音视频对话的合成、基准与训练
arXiv:2609.21465 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

OmniVChat-Studio,一种用于合成单轮与多轮音视频对话的多 Agent 数据引擎;以及 OmniVChat-RL,一种在 OmniVChat 中同时针对回复正确性、效率与风格设计奖励的强化学习奖励方案。OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues and OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat.

SteerDuplex: Steerable Duplex Speech Dialogue Models
SteerDuplex:可操纵的全双工语音对话模型
arXiv:2609.12623 多模态 方法 OA · 绿色 被引 2 · S2

提出 SteerDuplex,一种基于 Moshi 的全双工语音模型,在自然对话和针对指令遵循、语音表达、推理与双工交互的合成对话上进行了微调,并通过两阶段混合奖励强化学习改善时序与回复连贯性。SteerDuplex is introduced, a Moshi-based full-duplex speech model fine-tuned on natural conversations and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction, and two-stage reinforcement learning with hybrid rewards to improve timing and response continuity.

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
1% 的 token 已足够:论 On-Policy 蒸馏中的梯度估计。
arXiv:2609.24432 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

基于信噪分解与候选集近似的 IER(信息效率比),可基于 IER 及其与现有效用分数的组合进行 token 选择,同时保留采样的反向 KL 训练目标。An information-efficiency ratio (IER) based on a signal-to-noise decomposition and a candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective.

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies
CARE:面向 Vision-Language-Action 策略的经验引导原子级纠错执行。
arXiv:2609.24118 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CARE(Corrective Atomic Robotic Execution),一种通过从执行中遇到的失败进行学习以提升恢复能力的框架,并引入 Failure State Recovery Benchmark (FSR-Bench),用于评估在局部偏差与结构异常下从中间失败态恢复的表现。This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies.

Grounded Action Model: 3D Grounding as a Foundation for Robotics
Grounded Action Model: 3D Grounding as a Foundation for Robotics
arXiv:2609.23863 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Grounded Action Models (GAMs),一种基于 3D grounding 构建的机器人基础模型新范式,可自主运行,并作为底层控制器由高层规划器通过多种输入模态进行控制,支持长时序与依赖记忆的操作。Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding, are proposed, which can be run autonomously and serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation.

Streaming Video Editing with Easy Adaptation
Streaming Video Editing with Easy Adaptation
arXiv:2609.24788 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SVEET 框架,仅需在预训练的双向视频扩散模型上训练,即可支持高质量自回归流式视频编辑,并提出解耦训练方案,显式强制视频可控性与模型因果性优化方向之间的正交性。This paper proposes SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion and proposes a decoupled training scheme that explicitly enforces the orthogonality between the optimization directions of video controllability and model causality.

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
arXiv:2609.23796 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Mira-Scene,一种组合式 3D 场景重建框架,将稀疏姿态回归替换为密集有界对应恢复,并引入多模态扩散 Transformer 联合生成物体几何与 CCM,使用模态专精的专家流配合共享注意力与位置编码以促进几何-布局一致性。Mira-Scene is presented, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery and introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency.

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing
UltraTex:释放 2K 多视角扩散模型用于 3D 纹理生成
arXiv:2609.23169 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UltraTex,一种用于高分辨率多视图扩散 3D 纹理生成的高效端到端框架,并引入 Background Token Dropping(在 DiT 主干前移除背景 token)与 Block-Sparse Attention(降低前景序列上的注意力计算)。This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision
TAPe+ML:面向多任务计算机视觉的紧凑结构化表示
arXiv:2609.20869 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,将部分建模负担从网络参数转移至结构化输入表示,可在更低的数据、内存与算力需求下支持紧凑的多任务视觉系统。The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.

Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models
用于量子增强扩散语言模型的 Circuit Hypernetwork
arXiv:2609.24657 多模态 方法 被引 0 · S2

提出 HyperQ,在冻结的掩码扩散语言模型上添加 token 条件化的量子残差分支,支持 token 条件化电路发射,作为一种可处理的量子增强语言建模架构方法。HyperQ is introduced, which adds token-conditioned quantum residual branches to a frozen masked-diffusion language model, and support token-conditioned circuit emission as a tractable architectural approach to quantum-augmented language modelling.

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
脚本感知的 Mixture-of-Experts 统一多语言场景文本识别
arXiv:2609.24058 多模态 应用落地 OA · 绿色 被引 1 · S2

本文构建了 TextMuSS-10M,一个涵盖 10 种文字、229 种语言的大规模合成场景文本数据集,并提出了 ScriptMoE,一种具备文字感知能力的 Mixture-of-Experts (MoE) 架构。该架构在精度上达到最高,且比 per-language experts 更简单、比 VLM 更轻量,同时精度优于两者。This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.