研究库 论文知识库
Papers · organized/paper_cards

论文

443 张论文卡片 · 多模态

开放获取 全部 绿色 · 1640
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
QCell:重组与对齐细胞查询用于重叠实例分割
arXiv:2608.29253 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 QCell,一种新颖的基于查询的模型,用于在显微镜场景中去重叠细胞实例,在多个 benchmark 上优于 SOTA 方法,在 ISBI2014 上取得 +2.2 AP 和 +2.7 AJI。QCell is presented, a novel query-based model that de-overlaps cell instances in microscopy scenes and outperforms state-of-the-art methods across multiple benchmarks, achieving +2.2 AP and +2.7 AJI on ISBI2014.

A Common Measure of Communication for Speech Brain-Computer Interfaces
面向语音脑机接口的通用通信度量指标
arXiv:2609.02887 多模态 方法 被引 3 · S2

推导出 open-vocabulary mutual information (OVMI),一种衡量 decoder 相对于用户可能希望传达词汇的参考分布所传达信息的信息论量度,为语音 BCI 社区提供了一种原则性方法以比较异构系统、改进词汇设计并衡量领域进展。Deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate, provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
Motion-Omni:面向口语对话的端到端联合语音与全身动作生成
arXiv:2609.04250 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 Motion-Omni,一个端到端框架,其中的 spoken dialogue model 原生输出显式的 facial expression 以及手部、上半身和下半身运动,这些输出直接由生成语音的 hidden states 生成。Motion-Omni is presented, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech.

The Attention Triangle in Audio-Video Models
音视频模型中的注意力三角
arXiv:2609.03586 多模态 方法 OA · 绿色 被引 1 · S2

研究揭示沿音视频边的路由是双向的:音频可影响视频生成,视频也可影响音频生成;模型参数中编码的偏差是泄漏的主要来源之一。It is revealed that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation, and biases encoded in the model's parameters and emerges as a major contributor to leakage.

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
ShallowStream:先浅层索引再深层回答的流式视频理解
arXiv:2609.02780 多模态 应用落地 OA · 绿色 被引 2 · S2

ShallowStream 是一个利用 MLLM 浅层同时进行帧编码与检索索引构建的新框架,性能与当前最强流式方法相当,同时将单帧 prefill 延迟与 10 秒端到端延迟分别降低至多 52.1 倍和 11.9 倍。ShallowStream is a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building and achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively.

To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
于鲜活情境中观世界:统一的室内-室外城市世界生成
arXiv:2608.05879 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

HoloWorld 是首个在统一连贯的 3D 城市世界中同时支持室内与室外生成的框架,构建于持续更新的跨尺度世界上下文之上。HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world, built on a continuously updated cross-scale world context.

UniMate: One Unified Model to Animate Diverse Skeletons
UniMate:一个统一模型驱动多样化骨骼动画
arXiv:2609.05415 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UniMate,一个统一的 foundation model,可从绑定骨骼的 3D 资产与文本提示合成任意骨骼的关节运动,无需测试时优化或针对每个骨骼的重新训练,在质量、泛化性与效率上均超越 SOTA 基线。UniMate is presented, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining, and outperforms state-of-the-art baselines in quality, generalization, and efficiency.

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
一个编辑器,多种编辑:一个用于多样化视频编辑的统一免训练框架
arXiv:2609.04190 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

视频编辑涵盖多种编辑范式,但在单一统一框架中同时实现高质量的指令引导与主体引导编辑仍具挑战性。我们提出 EditVid,一个免训练框架,结合用于局部一致性的稀疏因果记忆、用于长程身份保持的基于对应关系的后注意力 token 注入,以及用于编辑局部性的软潜在融合。同一框架支持指令引导和参考引导的编辑,包括风格迁移、属性修改、对象插入、部分级编辑和主体替换。在Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On

The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
2026 PNPL 竞赛:LibriBrain100 中的词分类与高效跨被试泛化
arXiv:2609.03231 多模态 方法 OA · 绿色 被引 1 · S2

为推进任务课程聚焦于大规模词分类,本竞赛设置两条互补赛道:Deep 赛道面向同一被试内部的大规模词分类,目标是追求最佳性能;Broad 赛道面向跨被试泛化。Advancing the curriculum of tasks to focus on word classification to focus on word classification at scale, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation.

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
展开世界:在强化空间推理中分解 4D 属性
arXiv:2609.03729 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FactoSR,一个因子化强化学习框架,显式解读视觉投影所塌缩的维度,并指出强化显式的、因子化的 4D 一致性是将 VLM 演化为稳健、具有世界感知能力的推理器的关键一步。FactoSR is presented, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection, and suggests that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
FlowBalance:基于验证器的 On-Policy 推理经验自改进
arXiv:2609.03241 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

FlowBalance 是一种以验证器为锚定的自改进方法,学习完整响应上的归一化分布,在 Qwen3-4B 和 Qwen3-8B 上相对 FlowRL 提升了平均性能,同时改善了训练速度与稳定性,避免了直接 OPSD 响应长度坍缩,并在受控的 AIME24 诊断中表现出更高的正确策略多样性。FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses, improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.

What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
还有什么需要修复?探索对话生成制品中修订传播的成本有效 test-time compute
arXiv:2609.03254 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文为该设定引入新基准,并基于该基准评估了九种修订方法,包括序贯反思与并行采样变体,使用 gpt-oss-20b/120b、gpt-5.4-mini 以及 qwen3.5-9b/27b/122b 进行测试。A new benchmark for this setting is introduced, and nine revision methods are evaluated, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark.

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
蒸馏前先验证:用于 On-Policy 蒸馏的 Prompt 级教师门控
arXiv:2609.02998 多模态 方法 OA · 绿色 被引 4 · S2

本文提出 Teacher-Gated On-Policy Distillation(教师门控的在线策略蒸馏),其核心原则是在引入密集监督前以 prompt 级别验证教师可靠性;在全部六个单领域设定下优于 Vanilla OPD,并在多领域训练下于两种规模上取得更高的七项基准平均成绩。Teacher-Gated On-Policy Distillation is introduced, built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted, and outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training.

Unlocking Lossless Speedups in LLMs via Discrete Diffusion
通过离散扩散释放 LLM 的无损加速
arXiv:2609.04010 多模态 方法 OA · 绿色 被引 2 · S2

大语言模型(LLM)的成功很大程度上归功于 next-token prediction(NTP),但其自回归(AR)结构需要缓慢的串行 token 生成。为克服这一瓶颈,我们提出扩散增强 LLM,一类新模型,在使用扩散从该分布中并行采样多个 token 的同时定义 AR 模型分布。我们将这些模型的参数解耦为两组:AR 权重,使用标准 NTP 目标训练;轻量扩散权重,训练用于同时生成多个 token。扩散权重Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weight

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
EmbodiedSkills:用于编排、训练和部署 VLA Agent 的统一框架
arXiv:2609.01281 多模态 应用落地 OA · 绿色 被引 2 · S2

EmbodiedSkills 是一个统一框架,将每项技能决策视为执行提案:运行时先检查前置条件再执行,执行后再验证结果,为将底层 VLA 策略转化为闭环具身系统提供可训练且可检视的 Agent 层。EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward and provides a trainable and inspectable agent layer for turning low-level VLA policies into closed-loop embodied systems.

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
一个症状,三个杠杆:On-Policy Self-Distillation 的批判性综述
arXiv:2608.25936 多模态 综述 OA · 绿色 被引 0 · S2 + OpenAlex

本综述将崩溃视为受三处杠杆共同支配的症状:信号施加位置(即 token 加权方式)、教师所见信息(即特权信息的性质)、以及信号变化时机(即教师的动态性与引导衰减)。This review treats collapse as a symptom governed by three levers: where the signal is applied, that is, how tokens are weighted; what the teacher is shown, that is, the nature of the privileged information; and when the signal changes, that is, the teacher's dynamics and the decay of guidance.

Omni Interaction Agent Technical Report
Omni Interaction Agent 技术报告
arXiv:2609.08977 多模态 方法 OA · 绿色 被引 2 · S2

提出了 Gander,一个原生多模态双工交互模型,基于 MiniCPM-o 4.5 构建,并通过异步 Agent 循环进一步适配实时交互,同时开源其模型、代码和数据,以推动社区进一步研究与开发。Gander is presented, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop and is released together with its models, code, and data to facilitate further research and development in the community.

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Mask Forcing:通过双噪声掩码 rollout 改进自回归视频扩散蒸馏
arXiv:2609.09123 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Mask Forcing 是一种双噪声掩码展开策略,通过扰动自回归学生的自展开来缓解由 reverse-KL 模式寻求引发的模式坍缩,能以更高视觉质量高效改进多种自回归视频扩散蒸馏方法,且无需引入真实视频数据或额外后训练阶段。Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking, improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
AuK 技术报告:用于语音生成与编辑的开源基础模型
arXiv:2609.08936 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 AuK,一个开源基础模型,通过自然语言指令与音频上下文的统一接口整合语音生成与编辑,在零样本与指令控制的语音生成以及通用指令引导编辑上取得领先性能。AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context, is introduced and leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing is demonstrated.

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
TANGO:基于全身 Vision-Language-Action 模型的杂乱环境中人形机器人导航
arXiv:2609.09158 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 TANGO,首个面向语言条件人形机器人在杂乱环境中通行的全身视觉语言导航框架,在视觉语言导航任务中达到 SOTA,并在需要避障的困难场景中超越强模块化基线。This work introduces TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments, and demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation.

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
TransNormal-2:基于几何约束的 Rectified Flow 与边缘感知解码用于精确法向量估计
arXiv:2609.06665 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TransNormal-2 是基于 FLUX.2 的整流流框架,采用单步确定性推理,针对 VAE 解码器两侧的重建退化问题,施加 RGB 引导的残差修正以降低局部于边界的解码误差,且不自由改写粗预测。TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses VAE reconstruction degradation on both sides of the VAE decoder, and applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction.

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
RoboSPA:VLA 模型能否超越简单场景与短时序任务?
arXiv:2609.05324 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布大规模机器人操控数据集与基准,用于诊断 VLA 模型的具身推理能力,并将 RoboSPA 确立为面向更强、更可靠、更具泛化性具身 Agent 的挑战性诊断基准。A large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models, and establishes RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents.

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Marigold V2:重访用于单目深度估计的扩散 Transformer
arXiv:2609.08084 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

Marigold V2 在应用于其他稠密回归任务(如表面法向量估计与本征图像分解)时取得 SOTA 结果。Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.

SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
SynthGait-19K:用于步态参数估计的物理约束合成视频数据集
arXiv:2609.08108 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

发现空间步态参数对视觉域偏移更敏感,且更好的 HMR 重建并不一定带来下游步态估计的改进;提出 GaitXFormer,作为直接基于 RGB 的参考模型用于步态参数估计。It is found that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation, and GaitXFormer is introduced as a direct RGB reference model for estimating gait parameters.

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
RelightFormer:用于多视角物体重光照的前馈生成式 Transformer
arXiv:2609.07414 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一个前馈式生成 Transformer,用于直接的单视图与多视图图像重光照,完全绕过显式本征属性估计,并采用置换不变的位置编码对称处理无序多视图输入,避免序列偏差。This work introduces a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation and employs permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias.

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting
NOAH:学习完整患者旅程的纵向多模态时序感知表示与预测模型
arXiv:2609.09140 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Noah 是一个时间感知、任务无关的生成式 Transformer 模型,对完整多模态患者旅程进行表征与预测,是该领域首个真正整体化的生成模型,支持自回归预测,并具备可选的时间控制、零样本分类与反事实干预模拟能力。Noah is a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey, and is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation.

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
视频为何仍然如此昂贵?视频与音视频 LLM 推理效率机制综述
arXiv:2609.10355 多模态 综述 OA · 绿色 被引 0 · S2 + OpenAlex

本综述涵盖视觉与音视频 VideoLLM 的推理效率机制,这些机制报告了参数数量、每输入 FLOPs、延迟、内存或视觉/音频 token 数量的具体削减量,并按方法所作用的 pipeline 阶段加以组织。This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.

The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
语义瓶颈:利用语义表示实现非侵入式语音解码
arXiv:2609.10296 多模态 方法 OA · 绿色 被引 1 · S2

本文介绍 Brain2Semantics2Text,一种通过中间语义嵌入空间重建文本的方法,并阐述了该方法的核心原理、实现方式以及缓解学习可靠神经-语义映射挑战的策略。This work introduces Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space and describes the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping.

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
X-AuT:面向语音 LLM 的渐进式音频编码器压缩与跨尺度蒸馏
arXiv:2609.11412 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 X-AuT,一个渐进式框架,通过简短的行为探针选择层组合,并通过表征对齐、跨尺度蒸馏、调度式学生策略监督以及 LoRA 微调来恢复被剪枝的模型。X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning, is introduced.

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
SpatialBlock:通过合成积木堆叠问题增强 LVLM 的空间智能
arXiv:2609.07064 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

引入 SpatialBlock-15k,一个包含 15,000 个堆块问题的合成数据集,涵盖 3D 到 2D 投影、视角变换和结构组合,并提出受人类认知发展启发的新范式:通过结构化堆块操作任务学习基础空间技能。This work introduces SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination and proposes a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks.

TempCloze: Can Video-LLMs Identify the Missing Middle?
TempCloze:Video-LLM 能否识别缺失的中间片段?
arXiv:2609.01515 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

Video-LLM 的时间推理基准常以语言为中介,留有来自选项措辞、答案相关性或语言先验带来的语言捷径空间,因此提出 TempCloze,一个用于评估 Video-LLM 视觉时间推理能力的视频完形填空基准。Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors, so TempCloze is introduced, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
在统一多模态模型中将图像 tokenizer 视为视觉语言的研究
arXiv:2609.09143 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

构建一个受控的纯自回归测试床,在文本、图像、文生图(T2I)和图生文(I2T)预测的多模态持续预训练中跟踪任务特定验证损失,表明更好的重建并不一定带来更低的任务特定损失或更强的下游性能,且图像分词器的选择在联合优化下会影响文本建模。A controlled pure-autoregressive testbed is built and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction, showing that better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that image tokenizer choice can affect text modeling under joint optimization.

ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
ActReview:基于反驳引导训练数据与评分量规奖励的可操作同行评审生成
arXiv:2609.09076 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

引入 ActReview,一个由反驳引导的后训练框架,将论文特定的诊断连接到具体、有依据的修订计划,并引入 ActReview-Bench,一个包含 1,000 个实例的人工整理基准,用于评估诊断质量和修订实用性。This work introduces ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans, and introduces ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness.

SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image
SNAP3D:基于单张图像的物理合理可装配 3D 部件生成
arXiv:2609.13146 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种物理引导框架,用于提升单图像部件感知 3D 生成效果,使其几何结构物理相容且连接稳定,并在部件接触面引入参数化连接器。This work proposes a physics-guided framework for improving single-image part-aware 3D generation with physically compatible geometry and stable connections, and introduces parameterized connectors at their contact surfaces.

StepAudio 3 Gen Technical Report
StepAudio 3 Gen 技术报告
arXiv:2609.12945 多模态 方法 被引 0 · S2

本研究提出了 StepAudio 3 Gen,一个通用音频生成模型,在统一框架下支持零样本文本到语音合成(TTS)、声音设计、人声生成、音效、音乐、风格化语音以及多种音频类型的混合生成。This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
基于合成语音构建并评估固定音色的泰语 TTS
arXiv:2609.03502 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

最终得到的 82M 参数模型 Wayu-Paxa-TTS-Edge 实现了无需参考音频的设备端泰语 TTS,并在三个系统中取得了最低的停顿位置错误率与词内停顿率。The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.