研究发现,早期 OPD 训练动态均呈现一种规律的 *useful-transfer* 区间,其中留出准确率(即 *gold score*, $G$)随 $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$(学生初始化在 token 级反向 KL 散度的平方根)近似线性上升。It is found that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization.
论文
199 张论文卡片 · 工程化
本文提出 ATLAS(Aligned Transport of Latent Structure),一种在显式保持关系几何结构的同时校准全局潜空间分布的训练目标,通过一维 Wasserstein-2 传输进行 Wasserstein 嵌入匹配来校准其边缘分布。This work introduces Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution and uses Wasserstein embedding matching to calibrate its marginal through one-dimensional Wasserstein-2 transport.
实验表明,屏蔽高熵偏移位置相比朴素的 SD 提升了分布外泛化能力,由此得到的自蒸馏 judges 在所评估的主观子类别上比基于结果监督 RL 训练的 judges 高出 2-9 个百分点,同时在客观子类别上保持竞争力。Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD, and the resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
基于 Rec 的检索在所有划分上都提升了回访一致性;当上下文跨度足以覆盖每次返回的首次访问时,可再带来 24% 到 30% 的提升,但会牺牲一定图像质量;超过该跨度后,继续增加长度不再带来收益。Rec retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps.
本文首次系统梳理了深度学习研究迄今为止的激活函数应用趋势,将实践中的使用情况与文献中的研究成果进行对比。This paper will be the first, to compile the trends in AF applications in practice against the research results from literature, found in deep learning research to date.
在统计物理中,多元硬核模型描述一个粒子系统,每个粒子拥有各自的逸度。用图论语言表述,该模型的配分函数对应多元独立多项式,即独立多项式的多重仿射推广,定义为 $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$,其中 $\mathcal{I}(G)$ 表示 $[n]:=\{1,2,\dots,n\}$ 上图 $G$ 的所有独立集。我们证明对于 $[n]$ 上的每个简单图 $G$ 以及 $λ_1,\dots,λ_n\geq 0$,\[ Z_G(λ_1,\dots,In statistical physics, the multivariate hard-core model describes a system of particles, each of which receives its own fugacity. In graph-theoretic language, the partition function of the model translates to the multivariate independence polynomial, i.e., the multiaffine generalisation of the independence polynomial, defined by $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$, where $\mathcal{I}(G)$ denotes the set of all independent sets in a graph $G$ on $[n]:=\{1,2,\dots,n\}$. We prove that for every simple graph $G$ on $[n]$ and $λ_1,\dots,λ_n\geq 0$, \[ Z_G(λ_1,\dots,
预训练 Transformer 仅利用其深度的一小部分来跟踪上下文中的引用。十三个基础模型仅能可靠地跟随 1.4–3.6 行,额外的预训练循环收益甚微。在一个早期层上训练的 rank-8 LoRA 在所有模型权重冻结的情况下扩展了这一计算能力。Qwen3-8B 在 24 行链上的精确准确率从 15.5% 提升至 99%;更长训练的 LoRA 可达 50 行。Ouro-1.4B 经过四轮循环达到 60 行,八轮后至少达到 160 行。该 LoRA 启动了一场接力:程序行通过中间层的一段短距离传递其链身份。冻结的 head 逐层读取渐进式进展信号……Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressi
结果将控制器所学到的行为范围与该范围内请求的准确性分离开来,尽管连贯性下限并非对每种特性都成立。The results separate the behavioral range learned by a controller from the accuracy of requests within that range, although the coherence floor does not hold for every trait there.
本文用三种以不同方式施加长度压力的方法(固定生成预算、逐样本长度目标、组相对长度奖励)对多种模型进行微调,发现它们对忠实性和可监控性具有不同的影响。This work fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward, and finds that it affects faithfulness and monitorability differently.
本文首次系统地探索了不同 LLM 的有效选择与组织,以培育更公平的 LLM 回答,并展示了 CBM 显著优于独立基线。This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses and show CBM substantially outperforms standalone baselines.
JEPA-TTT 在测试时间内持续适配预训练动作条件 Joint-Embedding Predictive Architecture 世界模型的潜在动力学预测器,平均将自回归潜在预测误差降低 83%,并将规划性能相较冻结的 JEPA 世界模型提升 153%。JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time, reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model.
结果表明,面向框架的预训练、经过验证的 SFT 以及显式的能力保留评估,分别针对领域特定可执行代码生成中不同的失效模式。The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.
表格基础模型通过以标注训练样本为条件进行上下文学习 (ICL) 来预测。与分离训练与推理的传统模型不同,这些模型必须在每次前向传播中处理所有训练样本,导致每次预测成本高昂。限制训练样本数量可降低成本,但会显著降低性能。我们不丢弃上下文,而是提出激活对齐 (activation alignment),利用完整上下文来教会模型在仅看到子集时如何行为。这是通过训练一个……Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every forward pass, making each prediction expensive. Restricting the number of training examples reduces this cost but substantially degrades performance. Instead of discarding context, we propose activation alignment, a method that leverages the full context to teach a model how to behave when seeing only a subset. This is achieved by trainin
提出 LLM-as-Jev,一种保留架构的框架,直接从带括号数字标识符的下一 token 概率中提取校准决策,并同时给出无需训练的推理方案与一种微调目标——通过树分解的列表式损失优化候选选择,并使用 KL 散度惩罚将辅助预测锚定到基模型。This work presents LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers, and provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties.
强化学习(RL)是推动大型基础模型走向自我改进的核心训练范式。本报告介绍 MiMo-V2.6 系列,这是一个通过扩展 RL 算力来推动模型智能边界的 omni-modal 模型族。在 RL 之前,我们在广泛的跨模态语料上进行 mid-training 以提供充足的探索空间,并基于预训练的 hybrid-SWA 架构构建稳健的基础设施以支持后续规模化。我们沿三个维度扩展 RL 算力:(1) 更大的 batch 和更高的吞吐量,采用异步训练来持续消费Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes
CONFLUX 是一个面向胸部 CT 的潜在扩散模型:由 3D 变分自编码器压缩每个体数据,整流流 Transformer 在潜在空间中生成,并以分类器从生成体中恢复所请求病征的可靠性作为奖励。CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space, that rewards how reliably a classifier recovers the requested findings from each generated volume.
首个面向预训练胸部 X 光报告生成器的无需训练 best-of-N 采样方案,显式建模"既往-当前-过渡"纵向先验,全面优于随机选择。This work presents the first training-free best-of-N sampling scheme for pre-trained chest X-ray report generators that is explicitly aware of this longitudinal prior to current transition, and outperforms random selection across the board.
WARP:一个直接从已发布权重还原微调模型训练数据混合比例的框架,抽取几何特征并映射至各领域占比,可采用无参数 softmax 读出器,或基于合成混合训练的 MLP 投影器。WARP is introduced, a framework that recovers a fine-tuned model's training mixtures directly from its released weights and extracts geometric features and maps them to domain proportions using either a parameter-free softmax readout or an MLP projector trained on synthetic mixtures.
本文提出了 Graph-PRefLexOR,这是一族基于图结构的推理模型,使用 Group Relative Policy Optimization 进行微调,将推理过程组织为显式阶段,分别用于机理探索、图构建、模式提取和假设合成,确立了面向图结构的强化学习作为通向可解释 AI 系统的路径,可应用于材料设计及其他科学领域的科学假设生成。Graph-PRefLexOR is developed, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization to organize reasoning into explicit phases for mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis, establishing graph-native reinforcement learning as a pathway toward interpretable AI systems for scientific hypothesis generation in materials design and other scientific applications.
本文提出 Monotonic Inference Policy Update (MIPU),一种两步式 LLM 强化学习框架:构建采样器引用的候选更新,并使用推理侧差距代理选择性接受同步候选;实验表明 MIPU 提升了平均推理性能与训练稳定性。Monotonic Inference Policy Update (MIPU) is introduced, a two-step LLM RL framework that constructs sampler-referenced candidate updates and selectively accepts synchronized candidates using an inference-side gap proxy, and experiments show that MIPU improves average reasoning performance and training stability.
本文提出 MultiDepth-3k (MD-3k),一个用于衡量深度层偏好与多层空间关系准确率 (ML-SRA) 的稀疏双层序数基准;领先的深度基础模型在标准 RGB 输入下表现出不同的层偏好,表明同一分层几何可在不同模型中被差异化地解析。MultiDepth-3k (MD-3k), a sparse two-layer ordinal benchmark for measuring depth-layer preference and multi-layer spatial relationship accuracy (ML-SRA), is introduced and leading depth foundation models exhibit diverse layer preferences under standard RGB input, showing that the same layered geometry can be resolved differently across models.
水电隧道检测对基础设施完整性至关重要,但人工方式效率低下且具有危险性。本文提出 FLISP(Fast LiDAR-IMU Synchronized Path Planner),一种面向 UGV-UAV 协同检测的无地图规划框架。不同于传统基于地图的范式,FLISP 具有三项核心贡献:(1) 统一架构,由单套 UGV 搭载的 LiDAR-IMU 驱动两平台的同步路径生成;(2) 平台特定的求解器,采用增强型萤火虫算法用于 UGV 避障,以及动态迭代优化器用于 UAHydropower tunnel inspection is critical for infrastructure integrity yet remains inefficient and hazardous using manual methods. We propose FLISP (Fast LiDAR-IMU Synchronized Path Planner), a mapless planning framework for cooperative UGV-UAV inspection. Unlike traditional map-based paradigms, FLISP features three core contributions: (1) a unified architecture where a single UGV-mounted LiDAR-IMU suite drives synchronized path generation for both platforms; (2) platform-specific solvers utilizing an enhanced Firefly Algorithm for UGV obstacle avoidance and a dynamic iterative optimizer for UA
实验表明,LISA 不仅能持续加速训练收敛并提升最终合成结果,还能促使侧网络特征在条件建模中更加解耦,且几乎无额外训练成本,推理成本为零。Experiments demonstrated that LISA can not only consistently accelerate the training convergence and improve final synthetic results, but also encourage the side network's features to be more disentangled for conditional modeling with negligible additional training cost and zero extra inference cost.
本文表明强化学习(RL)后训练已具备实现有效 step-level 评分所需的要素,从而完全无需额外的奖励模型训练,并在通用随机 Markov 决策过程下推导出一种隐式 advantage,称为 progress advantage。This work shows that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether, and derives an implicit advantage under a general stochastic Markov decision process, which is term progress advantage.
本文提出 Randomized YaRN,一种通过将基于 YaRN 的位置外推与随机位置编码和长度课程相结合来提升长度泛化能力的训练方法,表明渐进式地将模型暴露于分布外位置分布是实现可泛化长上下文推理的有效方案。Randomized YaRN is proposed, a training method that improves length generalization by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum, and suggests that progressively exposing models to OOD positional distributions provides an effective recipe for generalizable long-context reasoning.
介绍 Stanford EDGAR Filings Dataset(SEFD),一个将 SEC filings 开放重建为版面保真 MultiMarkdown 的数据集,用于金融语言建模与评估;同时推出两个基于 SEFD 的基准:EDGAR-Forecast,用于评估模型知识截止后基于 filings 的数值预测;EDGAR-OCR,用于评估复杂金融表格的转录质量。The Stanford EDGAR Filings Dataset (SEFD), an open reconstruction of SEC filings into layout-faithful MultiMarkdown for financial language modeling and evaluation, is introduced and two SEFD-derived benchmarks are introduced: EDGAR-Forecast, which evaluates filing-grounded numerical forecasting after model knowledge cutoffs, and EDGAR-OCR, which evaluates transcription of complex financial tables.
结果表明,由 LLM 编排的多 Agent 系统可将传统 AutoML 扩展为可信、自适应且面向生产的 BDaaS 生命周期自动化。The results suggest that LLM-orchestrated multi-agent systems can extend conventional AutoML toward trustworthy, adaptive, and production-oriented BDaaS lifecycle automation.
本文提出 MedEasy,一个多 Agent 系统,通过患者对话、临床操作、决策提交、文档记录与反馈来组织虚拟患者练习,为使用案例特定标准连接情境化实践的 AI 辅助职业训练系统贡献了设计启示。MedEasy is presented, a multi-agent system that organizes virtual-patient practice through patient dialogue, clinical actions, decision submission, documentation, and feedback that contributes design implications for AI-supported professional training systems that use case-specific standards to connect situated practice.
CoRe(Context Relevance)被提出。该系统在一个大型短视频搜索引擎中每周重新部署,持续运行超过五个月,使用已部署的多模态相关性模型作为源,并以镜像生产融合代数的乘性比值形式来缩小仿真与生产之间的差距。CoRe (Context Relevance) is presented, such a system, redeployed weekly for over five months in a major short-video search engine, using the deployed multimodal relevance model as its source and a multiplicative ratio form mirroring the production fusion algebra to close the simulation-production gap.
提出 OmniTacTune,一种策略无关的真实世界 RL 流程,通过残差修正将触觉反馈适配到预训练视觉策略,并在多种接触丰富任务、视觉基础策略与触觉表征间实现泛化。OmniTacTune is introduced, a policy-agnostic real-world RL pipeline that adapts tactile feedback to pretrained visual policies through residual correction and generalizes across diverse contact-rich tasks, visual base policies, and tactile representations.
给出一个具体且有实证支撑的像素级 EO 基础模型扩展方案:训练大型 encoder,按下游性能筛选,再蒸馏为灵活的学生模型。A concrete, empirically grounded recipe for scaling pixel-wise EO foundation models: train large encoders, select by downstream performance, and distil into flexible student models is given.
借助强大的全景先验,Canvas360 构建了一个统一的上下文全景生成框架,通过 token 级拼接支持多样化下游任务,在任务覆盖范围与建模灵活性上均超越已有方法。Empowered by strong panoramic priors, Canvas360 enables a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility.
提出 PAST-TIDE,在官方排行榜上 Subtask A 取得 0.75 的 macro-F1,Subtask B 取得 0.74,表明对预训练模型仅做极少的架构改动即可在低资源场景下保持竞争力。PAST-TIDE is introduced, PAST-TIDE achieves macro-F1 scores of 0.75 for Subtask A and 0.74 for Subtask B on the official leaderboard, indicating that minimal architectural additions to a pre-trained model can remain competitive in low-resource settings.
通过将疾病特异性上下文整合到分子生成中,DrugGen-2 推动了 AI 辅助药物发现,为 de novo 设计和药物再利用提供了强大工具,可同时考虑疾病与分子靶点之间的复杂相互作用。By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.
本工作描述了如何在开源实现中满足使用有限项数的截断态向量来模拟 peaked circuits 的要求,并讨论了其性能与局限性。This work describes how the requirements to simulate peaked circuits using a truncated state vector with a limited number of terms were met in an open-source implementation, and discusses its performance and limitations.
论文提出 PanoWorld,通过固定朝向将相机轨迹简化为平移,并借助 Dense Panoramic Ray-Conditioning 与 Geometry-aware Memory Augmentation 同时支持当前动作建模与长程记忆。PanoWorld is proposed, which simplifies camera trajectories into translations via fixed headings for both current-action modeling and long-range memory through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation.