研究库 论文知识库
Papers · organized/paper_cards

论文

329 张论文卡片 · 多模态 · 方法

开放获取 全部 绿色 · 1640
GameWAM: A World Action Model for Video Games
GameWAM:面向视频游戏的世界动作模型
arXiv:2608.26200 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 GameWAM,据其所知是首个面向原生闭环游戏与 GUI 控制的 WAM,并发现 Low-Frequency Action Source Imprinting (LASI):在固定条件下,采样动作源的低频分量系统性地引导生成的粗粒度相机运动,揭示了生成式控制中的源敏感性失效模式。This work introduces GameWAM, to its knowledge the first WAM for native closed-loop gameplay and GUI control and uncovers Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control.

EditaLive! Unified Character Video Editing for Live Streaming
EditaLive! 面向直播的统一人物视频编辑
arXiv:2608.27123 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

传统视频编辑主要关注场景级内容,而直播更强调人物主体。然而,直接将现有视频编辑方法应用于以人为中心的直播仍具挑战,因为它们可能引入面部表情不一致,且通常依赖多个离线推理步骤,难以满足实时交互需求。我们提出 EditaLive,一个用于实时流式角色视频编辑的新型框架。具体而言,我们从预训练图像动画模型(Wan-Animate)出发,该模型天然解耦...Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decoupl

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
TacForcing:融合执行时间触觉反馈的流式动作生成
arXiv:2608.25798 多模态 方法 OA · 绿色 被引 1 · S2

提出 TacForcing,一种融合执行期触觉反馈的流式动作生成框架:按顺序生成动作块并保留未完成块的中间状态,并引入执行感知触觉注意力(EATA),将每次触觉更新的直接访问限制在下一个即将执行的块。TacForcing is introduced, a streaming action-generation framework incorporating execution-time tactile feedback that generates action blocks sequentially while preserving intermediate states of unfinished blocks and introduces Execution-Aware Tactile Attention (EATA), which restricts direct access to each tactile update to the next block scheduled for execution.

Luce: Relightable Gaussians for 3D Asset Generation
Luce:用于 3D 资产生成的可重光照高斯表示
arXiv:2608.23943 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Luce,一种三维表示方法,将几何与 PBR 材质统一到体素化的多模态高斯云中,分别使用专用高斯基元表示反照率、金属度-粗糙度与法线,在单图到三维生成任务上达到 SOTA。Luce is proposed, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for albedo, metallic-roughness, and surface normals to achieve state-of-the-art single-image-to-3D generation.

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
LayerRecall:面向视频生成长程一致性的状态条件记忆路由器。
arXiv:2608.28460 多模态 方法 OA · 绿色 被引 2 · S2

提出 LayerRecall,一种由当前状态条件化、按层选择的 memory router,仅从历史 K/V state 中检索相关状态并注入到 backbone 特定的 memory-sensitive 层中,其余层保留局部 attention;并提出 Cross-Horizon Prediction Matching (CHPM),借助特权的长上下文参考在预测空间中对有界 memory router 进行监督。This work introduces LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere, and proposes Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space.

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
J-Zero:从零数据出发的统一 Challenger--Solver--Judge 协同进化
arXiv:2608.26582 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Judge co-adaptation from Zero data (J-Zero),一个统一的 Challenger–Solver–Judge 自进化框架,支持在可验证与不可验证领域的自我提升,并识别出 Judge 协同进化是这一持续改进的关键驱动力。This work proposes Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both verifiable and unverifiable domains and identifies Judge co-adaptation as the key driver of this sustained improvement.

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
Ring Forcing:面向自回归视频扩散的精确长时记忆
arXiv:2608.26794 多模态 方法 OA · 绿色 被引 5 · S2

提出 Ring Forcing,一个自回归视频扩散框架,旨在稳健构建并精确利用长期记忆,并通过稀疏 RoPE 机制实现灵活、可扩展的记忆适配,同时充分利用预训练先验。Ring Forcing is presented, an autoregressive video diffusion framework designed to robustly construct and precisely utilize long-term memory and a sparse RoPE mechanism to enable flexible, scalable memory adaptation while fully exploiting pre-trained priors.

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
按意图行动:面向视觉-语言-动作模型的行为意图蒸馏
arXiv:2608.23478 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Intention Distillation (INDI),将行为级意图蒸馏到动作解码器中,并以目标依赖的方式组织下游预测;研究表明,动作解码器显式建模其生成行为的语义目标能够带来收益。Intention Distillation (INDI) is proposed, which distills behavior-level intent into the action decoder and organizes downstream predictions in an objective-dependent manner, and shows that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
LLM 研究论文:2026 年清单(1—5 月)— Sebastian Raschka
arXiv:2602.08071 多模态 方法 Open MIND OA · 绿色 被引 13 · S2

ViT-5 的设计与当代基础模型实践保持一致,可作为对 vanilla ViT 的直接替换升级方案,适用于 2020 年代中期的视觉骨干网络,并为生成建模提供更强大的骨干。With a design aligned with contemporary foundation-model practices, ViT-5 offers a simple drop-in upgrade over vanilla ViT for mid-2020s vision backbones and serves as a stronger backbone for generative modeling.

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
DreamX-Creator:在 2K 分辨率下实现原生音视频生成的民主化
arXiv:2608.31106 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过发布紧凑的 7B 生成器和 2K Refiner,该工作致力于使原生音视频生成平民化,并为统一音视频生成建模的未来研究提供可及的基础。By releasing the compact 7B generator and 2K Refiner, this work seeks to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.

Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models
Chat-Edit-3D++:基于大语言模型的交互式 3D 与 4D 场景编辑
arXiv:2608.29137 多模态 方法 OA · 绿色 被引 1 · S2

提出 Hash-Atlas 网络,将 3D 场景编辑重新表述为对 2D atlas 图像的操作,从而实现 2D 编辑与 3D 重建流程的解耦。The Hash-Atlas network is proposed, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes.

Unsupervised Learning of Video Representations using LSTMs
基于 LSTM 的视频表示无监督学习
arXiv:1502.04681 多模态 方法 OA · 绿色 被引 2770 · S2

本文利用 LSTM 网络学习视频序列表示,并通过在 UCF-101 与 HMDB-51 数据集上对人体动作识别这一监督学习任务进行微调来评估所学表示。This work uses Long Short Term Memory networks to learn representations of video sequences and evaluates the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
SpanCalib-VLM:视觉语言模型中校准的幻觉片段检测
arXiv:2608.29974 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SpanCalib-VLM,这是一种用于 SHROOM-Visions 共享任务的混合双系统,结合了多模态序列标注器(由 XLM-RoBERTa-Large 与 SigLIP 视觉编码器通过交叉注意力融合而成)与微调后的生成式 VLM(Qwen3.5-4B-SHROOM-SFT)。SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).

MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
MMMMM:用于研究多语言多模态虚假信息机制的统一分类法
arXiv:2608.29681 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

基于对数据的深入定性分析和既有理论研究,构建了一套新颖且全面的多模态错误信息分类法,并由此获得了关于社交媒体用户如何在实际场景中将图像与文本结合以传播错误信息的此前未被记录的洞见。A novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work is developed, which leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild.

RECAP-Forcing: Retaining Content Appearances for Long Video Generation
RECAP-Forcing:面向长视频生成的内容外观保持方法
arXiv:2608.26671 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 RECAP-Forcing,一种无需训练的推理方法,不增加任何可学习参数,在多个强基线上稳定提升视觉质量与语义保真度,并优于现有记忆方法。This work proposes RECAP-Forcing, a training-free inference method with no additional learnable parameters that consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
ZimaBlue:通过可扩展视频预训练演化出可泛化的通用世界动作模型
arXiv:2609.00188 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 ZimaBlue,一个可扩展的框架,用于从大规模视频中学习可泛化的 World Action Models (WAMs),并采用异步 Slow-Fast 双系统架构,使生成式 WAM 具备面向实时控制的实用性。This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Qwen-Drive-1.0:迈向自动驾驶视觉-语言基础模型的初步探索
arXiv:2609.00111 多模态 方法 OA · 绿色 被引 7 · S2

实验表明,该方法在大幅保留通用视觉-语言能力的同时,具备出色的 3D 感知与驾驶场景理解能力;在开环、伪闭环与闭环设定下的综合评估进一步显示其运动规划性能具有很强的竞争力。Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability and comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
学习结果变化的位置:面向多模态几何的可寻址信用推理
arXiv:2608.30457 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 credit-addressable reasoning:推理时暴露的语义单元同时定义学习阶段比较候选与分配 credit 的位置;并实例化为 Code-CoT,保留图示、将视觉关系表示为行可寻址的可执行代码,并将推理组织为类型化事件。This work introduces credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit, and instantiates Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events.

VibeVoice-ASR-Streaming Technical Report
VibeVoice-ASR-Streaming 技术报告
arXiv:2609.02812 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

VibeVoice-ASR-Streaming 是首批基于 LLM 的端到端流式说话人归属 ASR 方法之一,无需独立 diarization 阶段即可在语音到达时输出"谁说了什么"。VibeVoice-ASR-Streaming is one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR, allowing the model to produce''who said what''as speech arrives, without a separate diarization stage.

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
NeoMME:用于高效微调与推理的单塔多模态原生多语言基础编码器
arXiv:2609.01657 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NeoMME,一个 260M 与 800M 参数的多模态多语言双向编码器系列,可在单个双向 Transformer encoder 中处理多语言文本与原始图像 patch。This work introduces NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder.

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems
面向对话推荐系统的零数据引导的实证研究
arXiv:2504.15476 多模态 方法 OA · 绿色 被引 7 · S2

结果表明:领域驱动的合成数据一致优于零样本提示与朴素合成基线;主动选择相比随机采样提升了数据效率;元数据与协同过滤信号各自提升选择质量;在低资源场景下,合成数据可优于稀缺的真实对话,并进一步对真实对话形成补充。The results show that domain-grounded synthetic data consistently outperforms zero-shot prompting and naive synthetic baselines; active selection improves data efficiency over random sampling; metadata and collaborative filtering signals each improve selection quality; and, in low-resource settings, synthetic data can outperform scarce real dialogues while further complementing them.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Puffin-World:以原生 3D 世界状态扩展统一多模态模型
arXiv:2609.04196 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Puffin-World,一种统一的多模态架构,集成物理理解、空间仿真与 3D 世界生成重建,无需依赖外部离线模块,可支持需要多任务协同的交错式闭环应用。Puffin-World is proposed, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules and enables interleaved closed-loop applications requiring synergy across multiple tasks.

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
选择、压缩、再投入:长视频 MLLM 中视觉 token 分配的控制变量研究
arXiv:2609.03820 多模态 方法 被引 0 · S2

过程中发现 AKS baseline 自身存在实现 bug,且两个 harness 在相同预算下运行相同已发布规则存在 0.74 分的差距,这表明此类对比应在同一受控 harness 内进行,而非跨论文比较。Along the way, an implementation bug in the own AKS baseline and a 0.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.

QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
QCell:重组与对齐细胞查询用于重叠实例分割
arXiv:2608.29253 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 QCell,一种新颖的基于查询的模型,用于在显微镜场景中去重叠细胞实例,在多个 benchmark 上优于 SOTA 方法,在 ISBI2014 上取得 +2.2 AP 和 +2.7 AJI。QCell is presented, a novel query-based model that de-overlaps cell instances in microscopy scenes and outperforms state-of-the-art methods across multiple benchmarks, achieving +2.2 AP and +2.7 AJI on ISBI2014.

A Common Measure of Communication for Speech Brain-Computer Interfaces
面向语音脑机接口的通用通信度量指标
arXiv:2609.02887 多模态 方法 被引 3 · S2

推导出 open-vocabulary mutual information (OVMI),一种衡量 decoder 相对于用户可能希望传达词汇的参考分布所传达信息的信息论量度,为语音 BCI 社区提供了一种原则性方法以比较异构系统、改进词汇设计并衡量领域进展。Deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate, provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
Motion-Omni:面向口语对话的端到端联合语音与全身动作生成
arXiv:2609.04250 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 Motion-Omni,一个端到端框架,其中的 spoken dialogue model 原生输出显式的 facial expression 以及手部、上半身和下半身运动,这些输出直接由生成语音的 hidden states 生成。Motion-Omni is presented, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech.

The Attention Triangle in Audio-Video Models
音视频模型中的注意力三角
arXiv:2609.03586 多模态 方法 OA · 绿色 被引 1 · S2

研究揭示沿音视频边的路由是双向的:音频可影响视频生成,视频也可影响音频生成;模型参数中编码的偏差是泄漏的主要来源之一。It is revealed that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation, and biases encoded in the model's parameters and emerges as a major contributor to leakage.

To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
于鲜活情境中观世界:统一的室内-室外城市世界生成
arXiv:2608.05879 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

HoloWorld 是首个在统一连贯的 3D 城市世界中同时支持室内与室外生成的框架,构建于持续更新的跨尺度世界上下文之上。HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world, built on a continuously updated cross-scale world context.

UniMate: One Unified Model to Animate Diverse Skeletons
UniMate:一个统一模型驱动多样化骨骼动画
arXiv:2609.05415 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UniMate,一个统一的 foundation model,可从绑定骨骼的 3D 资产与文本提示合成任意骨骼的关节运动,无需测试时优化或针对每个骨骼的重新训练,在质量、泛化性与效率上均超越 SOTA 基线。UniMate is presented, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining, and outperforms state-of-the-art baselines in quality, generalization, and efficiency.

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
一个编辑器,多种编辑:一个用于多样化视频编辑的统一免训练框架
arXiv:2609.04190 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

视频编辑涵盖多种编辑范式,但在单一统一框架中同时实现高质量的指令引导与主体引导编辑仍具挑战性。我们提出 EditVid,一个免训练框架,结合用于局部一致性的稀疏因果记忆、用于长程身份保持的基于对应关系的后注意力 token 注入,以及用于编辑局部性的软潜在融合。同一框架支持指令引导和参考引导的编辑,包括风格迁移、属性修改、对象插入、部分级编辑和主体替换。在Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On

The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
2026 PNPL 竞赛:LibriBrain100 中的词分类与高效跨被试泛化
arXiv:2609.03231 多模态 方法 OA · 绿色 被引 1 · S2

为推进任务课程聚焦于大规模词分类,本竞赛设置两条互补赛道:Deep 赛道面向同一被试内部的大规模词分类,目标是追求最佳性能;Broad 赛道面向跨被试泛化。Advancing the curriculum of tasks to focus on word classification to focus on word classification at scale, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation.

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
展开世界:在强化空间推理中分解 4D 属性
arXiv:2609.03729 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FactoSR,一个因子化强化学习框架,显式解读视觉投影所塌缩的维度,并指出强化显式的、因子化的 4D 一致性是将 VLM 演化为稳健、具有世界感知能力的推理器的关键一步。FactoSR is presented, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection, and suggests that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
FlowBalance:基于验证器的 On-Policy 推理经验自改进
arXiv:2609.03241 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

FlowBalance 是一种以验证器为锚定的自改进方法,学习完整响应上的归一化分布,在 Qwen3-4B 和 Qwen3-8B 上相对 FlowRL 提升了平均性能,同时改善了训练速度与稳定性,避免了直接 OPSD 响应长度坍缩,并在受控的 AIME24 诊断中表现出更高的正确策略多样性。FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses, improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.

What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
还有什么需要修复?探索对话生成制品中修订传播的成本有效 test-time compute
arXiv:2609.03254 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文为该设定引入新基准,并基于该基准评估了九种修订方法,包括序贯反思与并行采样变体,使用 gpt-oss-20b/120b、gpt-5.4-mini 以及 qwen3.5-9b/27b/122b 进行测试。A new benchmark for this setting is introduced, and nine revision methods are evaluated, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark.

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
蒸馏前先验证:用于 On-Policy 蒸馏的 Prompt 级教师门控
arXiv:2609.02998 多模态 方法 OA · 绿色 被引 4 · S2

本文提出 Teacher-Gated On-Policy Distillation(教师门控的在线策略蒸馏),其核心原则是在引入密集监督前以 prompt 级别验证教师可靠性;在全部六个单领域设定下优于 Vanilla OPD,并在多领域训练下于两种规模上取得更高的七项基准平均成绩。Teacher-Gated On-Policy Distillation is introduced, built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted, and outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training.

Unlocking Lossless Speedups in LLMs via Discrete Diffusion
通过离散扩散释放 LLM 的无损加速
arXiv:2609.04010 多模态 方法 OA · 绿色 被引 2 · S2

大语言模型(LLM)的成功很大程度上归功于 next-token prediction(NTP),但其自回归(AR)结构需要缓慢的串行 token 生成。为克服这一瓶颈,我们提出扩散增强 LLM,一类新模型,在使用扩散从该分布中并行采样多个 token 的同时定义 AR 模型分布。我们将这些模型的参数解耦为两组:AR 权重,使用标准 NTP 目标训练;轻量扩散权重,训练用于同时生成多个 token。扩散权重Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weight