PLC-DPO 将每个偏好对的训练信号路由为 clean、flip 或 tie 三类,将噪声偏好学习从单纯过滤可疑样本重构为主动修正监督方向与强度。PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.
论文
199 张论文卡片 · 工程化
这是首个端到端证明:一个七人独立团队可以训练出在 agentic 网络安全能力上领先的开源权重模型,且全部三个 checkpoint 在相近参数规模下均排名第一。This is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability and all three checkpoints rank 1st among models at comparable parameter scales.
提出 ESRL,一种架构感知的框架,显式探索 MoE 模型的专家路由空间,将高置信度专家保留为锚点,并把随机路由限制在合理候选池内,从而保留可靠的计算路径。ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.
该工作提出了最小干预强化学习(Minimal Intervention Reinforcement Learning, MInTRL),通过在原本的 on-policy rollout 中引入稀疏的局部干预来扩展探索边界,确立了最小干预作为增强 on-policy RL 的有效范式。This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.
论文提出了 ModAR,这是首个在预测动作前以自回归方式对多种未来模态进行去噪的 WAM;研究发现 ModAR 的序列化生成优于现有 WAM 形式,并在所有评估数据规模下取得最高的平均成功率。ModAR is introduced, the first WAM to autoregressively denoise multiple future modalities before predicting actions, and it is found that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales.
Mind2Dialogue 框架提出了一个心理学引导的模拟器,在交互过程中保留个人特征并更新心智状态以生成连贯对话;通过对 Oracle 信息充分的回复进行训练,使模型在部署时无需直接访问用户心智状态即可提供帮助。The Mind2Dialogue framework proposes a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations, and trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment.
该工作表明,专用优化隐式地从此潜空间轨迹中进行选择,并建立了一种新的专用训练视角:当缺少 gold reasoning 时,调参选择直接控制传递给下游模型的潜在监督信号。This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.
我们推出 LimiX 家族新模型 LimiX-2,通过先前建立的 scaling laws 指导模型与数据规模扩展。LimiX-2 采用上下文机制网络 (CMNs) 范式,并以上下文条件掩码建模 (CCMM) 进行预训练。CMNs 将上下文学习的组织原则从以目标为中心的预测转向以机制为导向的联合建模。它并非围绕传统表格 PFN 的 p(y|x, D_context) 目标设计网络,而是围绕学习 p(x, y|D_context)——一种上下文依赖的表征We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the p(y mid x, D_{context}) objective of conventional tabular PFNs, it is designed around learning p(x, y mid D_{context}), a context-dependent representat
常用的并行方案留下了四种未被约束的并行维度,各自增长方式不同:随路由矩阵增长的专家调度、随 token 数乘词表规模增长的词表投影、随深度乘序列长度增长的梯度检查点边界,以及随参数量增长的优化器状态。Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.
该工作展示了一个紧凑的移动机械手,复用其手指同时完成运动与交互,无需独立的运动机构。This work demonstrates a compact mobile manipulator that reuses its fingers for locomotion and interaction, without a separate locomotion mechanism.
本文提出 When2Think,一种基于 RLVR 的后训练框架,用于实例自适应计算分配,既无需学习奖励模型,也无需学习 critic;离线参考缓存机制避免了策略更新阶段对参考模型的在线查询。This work proposes When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation that requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates.
本文提出 Srijika,一个为九种 Brahmic 文字(Devanagari、Tamil、Bengali、Telugu、Kannada、Malayalam、Gujarati、Gurmukhi、Odia)生成可安装 OpenType 字体的系统,并附带一份负面结果目录,覆盖参考引导重风格化中失败的 conditioning、目标函数选择和数据凸包限制。Srijika is presented, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia, and a negative-results catalogue covering failed conditioning, objective choices, and data-hull limits of reference-guided restyling.
提出 Calibrated On-Policy Distillation:通过正、负特权干预估计教师的自偏离区间,并仅保留超出该区间的部分以校准原始的 teacher–student 差异。Calibrated On-Policy Distillation is introduced, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region.
本文使用 Fast Iterative Shrinkage-Thresholding Algorithm 展开 CSC 优化,并将稀疏系数视为可微变量与网络参数联合学习;同时提出一种 label-free 的后训练策略,在固定主网络参数的情况下,根据被损坏输入自适应调整压缩强度。This work unfolds the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm and treats the sparsity coefficient as a differentiable variable jointly learned with the network parameters and introduces a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed.
本文探讨 Paxos 与 Raft 何者更优地解决分布式共识问题,通过以 Raft 的术语与实用抽象描述一个简化的 Paxos 算法,精确揭示两者的差异。This paper considers the question of which algorithm, Paxos or Raft, is the better solution to distributed consensus to determine exactly how they differ by describing a simplified Paxos algorithm using Raft's terminology and pragmatic abstractions.
RefineEdit 是一个基于 GRN 的 training-free prompt-to-prompt 图像编辑框架,将 bit routing 与两种稳定机制(adaptive spatial freezing 与 finite bit locking)相结合,使编辑证据能够随图像演化而被修正。RefineEdit is a training-free prompt-to-prompt image editing framework built on the GRN that combines bit routing with two stabilization mechanisms: adaptive spatial freezing and finite bit locking, allowing editing evidence to be revised as the image evolves.
论文证明,压缩效果的评估应采用按模态分层的下游能力保留度,而非单一精度指标,并为此建立了一套可复现的协议,用于对受监管的干细胞成像场景中的压缩基础模型进行审计。It is demonstrated that compression should be evaluated by modality-stratified downstream retention rather than single-number accuracy, and a reproducible protocol for auditing compressed foundation models in regulated stem-cell imaging is established.
提出 PARTS(Policy Adaptation with RL on Targeted Subtasks),一种真实世界子任务强化学习框架,能够将训练集中于瓶颈环节,同时以最小的人工干预推进训练 rollout。PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at bottlenecks while allowing training rollouts to proceed with minimal human intervention, is presented.
证明了对通用语言模型适配教育任务(涵盖问题求解、课程理解与教学辅助)的精选、能力均衡的监督数据的价值。The value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support is demonstrated.
提出 ACLArena,一个全面研究、分析与评估 Agent 持续学习(ACL)的框架,并提出新的 ACL 方案:将高质量轨迹的离线回放与多个由 RL 专精化的 LoRA 专家路由网络相结合,显著提升 Agent 跨多领域学习的能力。This work introduces ACLArena, a framework for comprehensively studying, analyzing, and evaluating Agent Continual Learning, and proposes a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains.
提出 Calibrated Clipping,一种动态方法,通过匹配下界裁剪分位数并相应重平衡上界,将 FP8 裁剪边界与高精度 BF16 分布对齐,消除熵激增并恢复与 BF16 基线可比的性能。Calibrated Clipping is proposed, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly, which eliminates entropy surges and restores performance comparable to the BF16 baseline.
提出 StableVQ,重新审视每个模块的合适学习目标,解决各模块独立训练以承担各自角色时产生的问题;其基于共享投影 codebook 构建,轻量且不引入可学习参数。StableVQ is proposed, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently, and Built on top of shared-projection codebooks, is lightweight and introduces no learnable parameters.
本文提出了 SEGMENT-SNAP,通过 part-handle coupling 融合几何与语义证据,在 Articulate3D Challenge 中获得第一名。This work presents SEGMENT-SNAP, which combines geometric and semantic evidence through part-handle coupling and achieved first place in the Articulate3D Challenge.
论文表明,在该高维空间中使用 flow matching 的标准速度预测要求模型拟合低维信号流形之外的正交噪声方向,导致优化效率低下,由此提出改用清晰数据参数化($\boldsymbol{x}_{0}$-prediction),使学习聚焦于底层信号流形。It is shown that using the standard velocity prediction in flow matching in this high-dimensional space requires the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient, and motivated using the clean data parameterization ($\boldsymbol{x}_{0}$-prediction) instead, which focuses learning on the underlying signal manifold.
PackLab 是一个用于开发、训练与评估闭环机器人装箱 MLLM 的综合框架,在不同物体集合与容器配置下均优于传统装箱启发式方法、经典强化学习方法以及通用 MLLM,展现了 MLLM 在长时任务机器人装箱中的潜力。PackLab is a comprehensive framework for developing, training, and evaluating MLLMs for closed-loop robotic bin packing that outperforms conventional packing heuristics, traditional reinforcement learning methods, and general-purpose MLLMs across object sets and container configurations, highlighting the potential of MLLMs for long-horizon robotic packing.
本文提出一个基于原理的、无需训练的框架,依次优化跨层权重配对与共享字典分解,并识别结构兼容的投影、学习一种能更好保留各层独立校准几何的共享表征。This work introduces a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations, and identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry.
主要发现是:多样且高质量的 SFT 奠定坚实的能力下限,难度过滤将 RL 提示维持在有效的学习区间内,而奖励可靠性为阶段排序提供了实用原则。The main findings are that diverse, high-quality SFT establishes a strong capability floor and difficulty filtering keeps RL prompts within a productive learning range, and reward reliability provides a practical principle for ordering stages.
本文给出证据,说明 superposition 是 Transformer 架构的内在属性,而非训练过程中涌现的结果,并证明通过轻量级 fine-tuning 可以在很大程度上恢复线性性。This work provides evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training, and demonstrates that linearity can be substantially restored through lightweight fine-tuning.
本文提出 CARD,一种通过渐进式细化实现有效个性化的层级框架:先按共享风格模式对用户聚类,再为各组学习专用的 LoRA adapter,从而在低资源场景下也能实现稳健的泛化与强劲的性能。This work presents CARD, a hierarchical framework that achieves effective personalization through progressive refinement that first clusters users according to shared stylistic patterns and learns group-specific LoRA adapters, enabling robust generalization and strong low-resource performance.
本文提出 ExpertAlign,一种无需域标签或独立 routing 模型的框架,可在每个 token 上对全体 teacher 池进行监督路由;研究表明 token 级路由能够利用跨域互补监督,并减少对 prompt 级域指派的单一依赖。This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment.
本文识别出两种强化水印证据对载体图像依赖性、抑制残余可迁移性的机制,并提出 CoverLock,一种即插即用策略,可在不重新设计架构的前提下强化现有水印系统的图像依赖性。This work identifies two mechanisms that strengthen the dependence of watermark evidence on the cover image, thereby suppressing the residual transferability, and introduces CoverLock, a plug-and-play strategy for existing watermarking systems that strengthens such image dependence without architectural redesign.
Nereus 是一种成本感知的运行时,将 RL 后训练任务适配为高效执行计划,并基于内存可行的全局计划执行状态转移,使用与运行任务校准的成本模型来接纳转移。Nereus is a cost-aware runtime that adapts RL post-training jobs into efficient execution plans and executes a transition using a memory-feasible global plan and admits the transition using a cost model calibrated against the running job.
注意力往往集中在上下文中一小部分 token 上,但每个 query 关注的关键子集各不相同。为利用这种动态结构,我们提出 SANTA++,一种免训练的随机注意力方法,通过代表性 key 进行内存高效的选择,无需扫描整个 KV cache。缓存的 key 被组织成若干 team,query 对每个 team 中的代表性 key 打分以决定采样哪些 team。我们在采样得到的 team 内计算精确的注意力分数,并通过其采样概率的倒数对各 team 的贡献进行重新加权。Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusio
结果表明,推理路径多样性可作为筛选 SFT 数据的实用准则,能更好地为 RL 准备模型,并据此提出一种轻量级、基于规则的指纹方法用于筛选。These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL, and propose a lightweight, rule-based fingerprint to select for it.
提出 Chinese-Jev,一个 System One 模型,通过统一的数据处理与训练流水线弥合预训练分布与下游中文场景间的差距,在各专业领域上达到 Jev 平均准确率的 92%。Chinese-Jev is introduced, a System One model that addresses the gap between the pre-training distribution and downstream Chinese-language scenarios through a unified data processing and training pipeline, and achieves 92% of Jev's average accuracy across specialized domains.
本文推导了对应的反向传播规则(包括校准项的导数),并在模型、数据、优化器均一致的预训练实验中对不同选择进行了比较。This work derives the corresponding backward rules, including calibration derivatives, and compares these choices in pretraining experiments matched on model, data, and optimizer, and compares these choices in pretraining experiments matched on model, data, and optimizer.