研究库 论文知识库
Papers · organized/paper_cards

论文

1094 张论文卡片 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
ModaLens:报告条件医学 VLM 中图像敏感性的度量
arXiv:2609.15635 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ModaLens 是一种配对图像交换审计,用于衡量报告可用性如何改变图像敏感性,并限制对视觉正确性的结论;该方向在另外两个模型系列中得到复现。ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity and limits conclusions about visual correctness; the direction replicates in two further model lineages.

2. 分布式向量数据库 Qdrant 在 HPC 上的性能(arXiv 2509.12384,2025-09,持续更新)
arXiv:2509.12384 数据与向量库 方法 OA · 绿色 被引 10 · S2

本文在 Argonne Leadership Computing Facility 的 Polaris 超级计算机上对分布式向量数据库性能进行了实证研究,选取 Qdrant 评估在最多 32 个 worker 下的插入、索引构建与 query latency。This work presents an empirical study of distributed vector database performance on the Polaris supercomputer in the Argonne Leadership Computing Facility, and selects Qdrant to evaluate insertion, index construction, and query latency with up to 32 workers.

Expert-Space Exploration in MoE Reinforcement Learning
MoE 强化学习中的专家空间探索
arXiv:2609.13058 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ESRL,一种架构感知的框架,显式探索 MoE 模型的专家路由空间,将高置信度专家保留为锚点,并把随机路由限制在合理候选池内,从而保留可靠的计算路径。ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.

MInTRL: Off-policy Intervention can boost On-policy RL
MInTRL:Off-policy 干预可增强 On-policy RL
arXiv:2609.12419 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了最小干预强化学习(Minimal Intervention Reinforcement Learning, MInTRL),通过在原本的 on-policy rollout 中引入稀疏的局部干预来扩展探索边界,确立了最小干预作为增强 on-policy RL 的有效范式。This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Dynin-Robotics:全模态统一扩散视觉-语言-动作模型
arXiv:2609.13053 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Dynin-Robotics 在 LIBERO 和零样本 LIBERO-Plus 上取得了具有竞争力的性能,并在 Franka Research 3 机器人的四种操作条件下达到了 78.4% 的平均成功率。Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.

CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering
CiteGuard-RAG:一种以验证为中心的、基于证据的问答 AI 系统
arXiv:2609.15830 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

外部评估显示,尽管引用有效性保持稳健,但在领域偏移下证据利用、片段对齐与拒答校准变得更加困难,表明可信的 RAG 系统需要在检索与最终答案交付之间进行显式验证。External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift, indicating that trustworthy RAG systems require explicit validation between retrieval and final answer delivery.

Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
驾驭稀疏证据:通过显式上下文选择与整合的 Agentic Visual RAG
arXiv:2609.15800 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了 SCoRE(Selection and Consolidation for Robust Evidence),一个用于显式证据选择与整合的统一 agent 循环,将最终推理与探索式试错解耦,并通过索引化的声明-图像关联确保严格的视觉锚定。This work proposes SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation, which decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages.

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation
ORDER:面向检索增强生成的任务条件化路由。
arXiv:2609.17012 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

检索增强生成(RAG)管线通常依赖在预处理阶段确定的固定索引与检索配置。这种一刀切的设计难以适配领域专家场景,因为异构查询需要不同的分块粒度、元数据约束与来源选择策略。因此,针对某一类查询有效的配置,往往在其他类查询上表现欠佳。本文提出 ORDER(Optimal Routing for Dynamic Evidence Retrieval),一种查询条件化的 RAG 框架,可联合自适应地调整索引与……Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and ret

Modality-Autoregressive World-Action Models
模态自回归世界-动作模型
arXiv:2609.17524 工程化 方法 OA · 绿色 被引 1 · S2

论文提出了 ModAR,这是首个在预测动作前以自回归方式对多种未来模态进行去噪的 WAM;研究发现 ModAR 的序列化生成优于现有 WAM 形式,并在所有评估数据规模下取得最高的平均成功率。ModAR is introduced, the first WAM to autoregressively denoise multiple future modalities before predicting actions, and it is found that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales.

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
Mind2Dialogue:通过模拟用户心理状态训练具备人类感知能力的语言模型
arXiv:2609.15972 工程化 方法 OA · 绿色 被引 1 · S2

Mind2Dialogue 框架提出了一个心理学引导的模拟器,在交互过程中保留个人特征并更新心智状态以生成连贯对话;通过对 Oracle 信息充分的回复进行训练,使模型在部署时无需直接访问用户心智状态即可提供帮助。The Mind2Dialogue framework proposes a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations, and trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment.

Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
无推理轨迹的专家模型训练:面向领域专家蒸馏
arXiv:2609.13770 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作表明,专用优化隐式地从此潜空间轨迹中进行选择,并建立了一种新的专用训练视角:当缺少 gold reasoning 时,调参选择直接控制传递给下游模型的潜在监督信号。This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
人类构建的最后一个AI:迈向真正的递归自我改进
arXiv:2609.11873 Agent 智能体 方法 OA · 绿色 被引 4 · S2

该工作利用 Headroom-Closed Index(HCI)揭示现有 LLM 的问题,并提出 RSI 概念及其发展路线图:从改进执行自主性、改进策略自主性、经验获取自主性、环境适应自主性,到递归元改进。This work uses the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, and introduces the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement.

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
HarnessVLN:通过智能体框架统一免训练的具身导航
arXiv:2609.15195 评测基准 方法 OA · 绿色 被引 4 · S2

本工作提出 HarnessVLN,一个零样本、无需训练的框架:通过共享的 Agent Harness 统一指令跟随与物体目标导航,并展示了其在真实世界中两类导航任务上的适用性This work introduces HarnessVLN, a zero-shot, training-free framework that unifies instruction-following and object-goal navigation through a shared Agent Harness and demonstrates its applicability to both navigation tasks in real-world environments.

2. DualPath:打破 Agentic LLM 推理的存储带宽瓶颈
arXiv:2602.21548 LLM 基础设施 方法 Open MIND OA · 绿色 被引 20 · S2

本文提出 DualPath,一种 inference 系统,通过引入 dual-path KV-Cache loading 打破瓶颈,并实现一条新的 storage-to-decode 路径:KV-Cache 先加载到 decode engine,再通过 compute network 上的 RDMA 高效转发至 prefill engine。DualPath is presented, an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading and enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network.

StepAudio 3 Realtime Technical Report
StepAudio 3 Realtime 技术报告
arXiv:2609.14005 多模态 方法 OA · 绿色 被引 3 · S2

StepAudio 3 Realtime 是一个围绕持续 listen-converse-think-act 循环组织的音频-语言基础模型,在实时语音输出的同时达到与专用推理模型相当的对话与推理性能,并通过 Think-While-Speaking 机制化解深度思考与延迟之间的矛盾。StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop, achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time, and resolves the tension between deep deliberation and latency via Think-While-Speaking.

StepAudio 3 Music Technical Report
StepAudio 3 Music 技术报告
arXiv:2609.16034 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了 StepAudio 3 Music,一个支持显式音乐规划与开放域文本控制生成的大规模长篇音乐生成模型,在所评估系统中取得最高的 AudioBox Content Enjoyment、Content Usefulness 与 Production Quality 分数,以及最高的 MuQ-MuLan 相似度。This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems.

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
FLAT:将图像和文本重采样为 1D 变长对齐跨模态 token 用于检索与生成
arXiv:2609.16591 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文重新审视联合多模态表示学习与生成,旨在产生可直接被生成式解码器使用的线性可插值嵌入,并确保其表示同时充当判别性语义描述符和生成条件。This work revisits joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders and ensures its representations function as both discriminative semantic descriptors and generative conditions.

LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
LimiX-2:迈向通用结构化数据智能的上下文机制网络
arXiv:2609.17488 工程化 方法 OA · 绿色 被引 4 · S2

我们推出 LimiX 家族新模型 LimiX-2,通过先前建立的 scaling laws 指导模型与数据规模扩展。LimiX-2 采用上下文机制网络 (CMNs) 范式,并以上下文条件掩码建模 (CCMM) 进行预训练。CMNs 将上下文学习的组织原则从以目标为中心的预测转向以机制为导向的联合建模。它并非围绕传统表格 PFN 的 p(y|x, D_context) 目标设计网络,而是围绕学习 p(x, y|D_context)——一种上下文依赖的表征We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the p(y mid x, D_{context}) objective of conventional tabular PFNs, it is designed around learning p(x, y mid D_{context}), a context-dependent representat

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
OmniHarness:通过符号化策略学习实现可泛化的视觉生成
arXiv:2609.16057 评测基准 方法 OA · 绿色 被引 1 · S2

本文提出 OmniHarness,一个通过符号策略学习实现可泛化视觉生成的框架,将已验证的执行抽象为视觉生成任务族的符号策略,捕获共享流程和适用条件,同时去除实例特定的输入。OmniHarness is introduced, a framework for generalizable visual generation via symbolic policy learning that abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs.

Convergent Emergence of In-Context Learning Across Modalities
跨模态上下文学习的会聚涌现
arXiv:2609.14011 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个受控的跨模态框架,在多种模态下实例化相同的任务套件以检验 Convergent Emergence Hypothesis,结果显示配对映射的 ICL 在六种模态中出现,超越受控基线,并在其中五种模态上呈现相关的逐任务效应。A controlled cross-modality framework that instantiates the same task suite in a variety of modalities to test the Convergent Emergence Hypothesis and shows that paired-mapping ICL emerges across six modalities, surpasses controlled baselines, and has correlated per-task effects across five of them.

RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
RelateAnything: 任意输入的实时开放词汇关系预测
arXiv:2609.12552 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 RelateAnything,一个 53M 参数的模型,输入一张图像和来自任意来源的区域,针对推理时以字符串形式提供的谓词词汇返回带分数的关系。This work presents RelateAnything, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings.

Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
广义 Agent 迭代:迭代式策略改进与递归自改进的统一形式化框架
arXiv:2609.13406 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Generalized Agent Iteration (GAI),一个形式化框架,将迭代策略改进和 RSI 描述为同一学习范式的两种情形,基于经典理论,使现有系统可比较,并为分析和设计新系统提供原则性基础。This paper proposes Generalized Agent Iteration (GAI), a formal framework that describes iterative policy improvement and RSI as two cases of a single learning paradigm that rests on the classical account, makes existing systems comparable, and provides a principled basis for analyzing and designing new ones.

Agora: Git as Shared Memory for Collective AutoResearch
Agora: 以 Git 作为集体 AutoResearch 的共享记忆
arXiv:2609.18094 Agent 智能体 方法 OA · 绿色 被引 2 · S2

报告了一次持续近 12 天的运行:13 个 LLM worker 在无任务分配、无中央规划器的情况下,使用 Agora 解决了一个权重迁移问题,并记录了 agent 如何复用与验证共享工作。A run of nearly 12 days is reported in which 13 language-model workers, with no assigned tasks or central planner, used Agora to solve a weight-transfer problem and documents how agents reused and verified shared work.

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
重新思考 PPO 中的 Critic 学习:理解与缓解 Value Flattening
arXiv:2609.18708 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文识别出 Value Flattening 是标准 PPO 中 critic 学习的一种重要但被忽视的失效模式,并提出一种简单稀疏监督策略可以缓解该问题;引入 SParse Proximal Policy Optimization,在每个响应中仅对少数间隔良好的状态施加 value loss,以同时缓解两种效应。Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both effects.

A Zeroth-Order Paradigm for LLM Preference Alignment
一种面向 LLM 偏好对齐的零阶范式
arXiv:2609.19144 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出并分析 Comparison-based Preference Optimization (ComPO),一种基于比较预言机 (oracle) 的零阶对齐方法,并在光滑性、梯度稀疏性以及 oracle 与潜在目标相容的条件下,为其基础离线方案建立了收敛性保证。This paper proposes and analyzes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles, and establishes a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective.

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
Zing-0.5: 迈向具备实时联合动作与文本控制的可玩世界
arXiv:2609.17909 多模态 方法 OA · 绿色 被引 2 · S2

我们提出 Zing-0.5,一个 5B 自回归世界模型,专注于可玩性:用户可以探索生成的世界、影响事件演进,并通过键盘与在线文本的联合控制对反馈做出响应。我们的方法汇聚了三项技术贡献:(1) 统一的动作与文本条件建模,将感知幅度的键盘输入与时序对齐的文本指令、以及联合标注的视频结合,在同一序列中学习导航与事件控制;(2) 面向增量生成的事件尺度监督,使用分段级教师...We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trai

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
HypoEvolve:遗传算法赋能多智能体 LLM 发现科学假设
arXiv:2609.15938 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出一种代际遗传算法,用以协调专门的 LLM Agent,整合机制性论证、重新审视假设并评估证据与可检验性,推动了自主科学的愿景——AI 研究团队实现超越单个模型的发现能力。This work proposes a generational genetic algorithm to coordinate specialized large language model agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability that advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.

SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
SpectralShift:通过谱重参数化实现 Gated DeltaNet 高效的上下文窗口扩展
arXiv:2609.14320 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SpectralShift,一种用于 GDN 长上下文持续预训练的谱重参数化方法,通过重参数化 alpha 投影的初始化以重塑衰减谱,增强慢传播能力,并进一步引入 alpha 投影的学习率缩放以促进长上下文训练。SpectralShift is proposed, a spectral reparameterization approach for long-context continual pretraining of GDNs that reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training.

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
PANORAMA:基于掩码提议选择的全景接地描述生成
arXiv:2609.19143 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PanoCaps,一个由全景分割数据集构建的人工标注 benchmark,并提出 PANORAMA,一个将预训练 segmenter 条件化于上下文化短语表示以获得候选 mask、并学习选择与每个短语对应的 mask 的 VLM。This work introduces PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets, and introduces PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase.

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
内存墙的另一半:通过训练路由预测从 SSD 服务 35B MoE
arXiv:2609.18063 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Edge0,一个流式 MoE 推理引擎,通过 prerouter 缩小差距:每层 head 提前一个 token 预测下一层的 routing,并将该预测直接作为 routing 使用,使分阶段 expert 集合等于路由集合,无任何丢弃。Edge0 is presented, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped.

CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
CERA-MoA:与持续学习 LLM Agent 协同进化的路由机制
arXiv:2609.18779 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文设计了一种预测式熟悉度估计器,利用中间层隐藏状态评估 Agent 间的语义能力,避免完整 rollout 的开销,并在任务性能和效率之间实现权衡。A predictive familiarity estimator that leverages mid-layer hidden states to evaluate semantic competence among agents, avoiding the overhead of full rollouts and achieving a trade-off between task performance and efficiency is designed.

Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
Fathom:面向卸载 KV 缓存稀疏解码的每查询读取深度
arXiv:2609.17652 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Fathom,一种 key scan,其中每个查询决定读取每个 key 通道的比特数;在真实的 coding agent 会话中,Fathom 以 92 bit 达到最准确的 136 bit scan 的步骤一致性。Fathom is presented, a key scan in which each query decides how many bits of each key channel to read, and on real coding-agent sessions Fathom reaches the step agreement of the most accurate 136-bit scan at 92 bits.

Assessing nnU-Net Generalization across Brain Tumor Populations in BraTS-GoAT 2026
在 BraTS-GoAT 2026 中评估 nnU-Net 跨脑肿瘤人群的泛化能力

BraTS-GoAT 在异质人群上使用传统 3D nnU-Net 对 1,351 个标注案例进行肿瘤分割评估,采用五折交叉验证,每折 1,000 epochs,并应用 test-time mirroring。BraTS-GoAT evaluates tumor segmentation across heterogeneous populations using a conventional 3D nnU-Net on 1,351 labeled cases using five-fold cross-validation and 1,000 epochs per fold and applied test-time mirroring.

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
长上下文 Mixture-of-Experts 训练中每个内存峰值的平整化
arXiv:2609.14306 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

常用的并行方案留下了四种未被约束的并行维度,各自增长方式不同:随路由矩阵增长的专家调度、随 token 数乘词表规模增长的词表投影、随深度乘序列长度增长的梯度检查点边界,以及随参数量增长的优化器状态。Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
RetireOPD:用于智能体强化学习的自退役 On-Policy 蒸馏
arXiv:2609.20784 Agent 智能体 方法 OA · 绿色 被引 1 · S2

该工作提出 RetireOPD(Self-Retiring On-Policy Distillation),先用环境奖励优化一个解耦的、技能条件化的教师模型,再联合 RL 与 OPD 训练一个无技能依赖的学生模型。This work proposes RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
WeVisDoc:从覆盖到能力的鲁棒端到端文档解析
arXiv:2609.20423 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 WeVisDoc,一个面向鲁棒端到端文档解析的两阶段数据驱动框架,通过异构数据与保持结构的退化合成来扩展语义、结构与外观覆盖度,并指导有针对性的数据构建与目标 token 预算的重分配。WeVisDoc is presented, a two-stage data-centric framework for robust end-to-end document parsing that broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis and guides targeted data construction and reallocation of the target-token budget.