我们推出 LimiX 家族新模型 LimiX-2,通过先前建立的 scaling laws 指导模型与数据规模扩展。LimiX-2 采用上下文机制网络 (CMNs) 范式,并以上下文条件掩码建模 (CCMM) 进行预训练。CMNs 将上下文学习的组织原则从以目标为中心的预测转向以机制为导向的联合建模。它并非围绕传统表格 PFN 的 p(y|x, D_context) 目标设计网络,而是围绕学习 p(x, y|D_context)——一种上下文依赖的表征We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the p(y mid x, D_{context}) objective of conventional tabular PFNs, it is designed around learning p(x, y mid D_{context}), a context-dependent representat
论文
1173 张论文卡片 · 方法
本文提出 OmniHarness,一个通过符号策略学习实现可泛化视觉生成的框架,将已验证的执行抽象为视觉生成任务族的符号策略,捕获共享流程和适用条件,同时去除实例特定的输入。OmniHarness is introduced, a framework for generalizable visual generation via symbolic policy learning that abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs.
本文提出一个受控的跨模态框架,在多种模态下实例化相同的任务套件以检验 Convergent Emergence Hypothesis,结果显示配对映射的 ICL 在六种模态中出现,超越受控基线,并在其中五种模态上呈现相关的逐任务效应。A controlled cross-modality framework that instantiates the same task suite in a variety of modalities to test the Convergent Emergence Hypothesis and shows that paired-mapping ICL emerges across six modalities, surpasses controlled baselines, and has correlated per-task effects across five of them.
本文提出 RelateAnything,一个 53M 参数的模型,输入一张图像和来自任意来源的区域,针对推理时以字符串形式提供的谓词词汇返回带分数的关系。This work presents RelateAnything, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings.
本文提出 Generalized Agent Iteration (GAI),一个形式化框架,将迭代策略改进和 RSI 描述为同一学习范式的两种情形,基于经典理论,使现有系统可比较,并为分析和设计新系统提供原则性基础。This paper proposes Generalized Agent Iteration (GAI), a formal framework that describes iterative policy improvement and RSI as two cases of a single learning paradigm that rests on the classical account, makes existing systems comparable, and provides a principled basis for analyzing and designing new ones.
报告了一次持续近 12 天的运行:13 个 LLM worker 在无任务分配、无中央规划器的情况下,使用 Agora 解决了一个权重迁移问题,并记录了 agent 如何复用与验证共享工作。A run of nearly 12 days is reported in which 13 language-model workers, with no assigned tasks or central planner, used Agora to solve a weight-transfer problem and documents how agents reused and verified shared work.
本文识别出 Value Flattening 是标准 PPO 中 critic 学习的一种重要但被忽视的失效模式,并提出一种简单稀疏监督策略可以缓解该问题;引入 SParse Proximal Policy Optimization,在每个响应中仅对少数间隔良好的状态施加 value loss,以同时缓解两种效应。Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both effects.
本文提出并分析 Comparison-based Preference Optimization (ComPO),一种基于比较预言机 (oracle) 的零阶对齐方法,并在光滑性、梯度稀疏性以及 oracle 与潜在目标相容的条件下,为其基础离线方案建立了收敛性保证。This paper proposes and analyzes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles, and establishes a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective.
我们提出 Zing-0.5,一个 5B 自回归世界模型,专注于可玩性:用户可以探索生成的世界、影响事件演进,并通过键盘与在线文本的联合控制对反馈做出响应。我们的方法汇聚了三项技术贡献:(1) 统一的动作与文本条件建模,将感知幅度的键盘输入与时序对齐的文本指令、以及联合标注的视频结合,在同一序列中学习导航与事件控制;(2) 面向增量生成的事件尺度监督,使用分段级教师...We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trai
本文提出一种代际遗传算法,用以协调专门的 LLM Agent,整合机制性论证、重新审视假设并评估证据与可检验性,推动了自主科学的愿景——AI 研究团队实现超越单个模型的发现能力。This work proposes a generational genetic algorithm to coordinate specialized large language model agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability that advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.
本文提出 SpectralShift,一种用于 GDN 长上下文持续预训练的谱重参数化方法,通过重参数化 alpha 投影的初始化以重塑衰减谱,增强慢传播能力,并进一步引入 alpha 投影的学习率缩放以促进长上下文训练。SpectralShift is proposed, a spectral reparameterization approach for long-context continual pretraining of GDNs that reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training.
本文提出 PanoCaps,一个由全景分割数据集构建的人工标注 benchmark,并提出 PANORAMA,一个将预训练 segmenter 条件化于上下文化短语表示以获得候选 mask、并学习选择与每个短语对应的 mask 的 VLM。This work introduces PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets, and introduces PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase.
本文提出 Edge0,一个流式 MoE 推理引擎,通过 prerouter 缩小差距:每层 head 提前一个 token 预测下一层的 routing,并将该预测直接作为 routing 使用,使分阶段 expert 集合等于路由集合,无任何丢弃。Edge0 is presented, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped.
本文设计了一种预测式熟悉度估计器,利用中间层隐藏状态评估 Agent 间的语义能力,避免完整 rollout 的开销,并在任务性能和效率之间实现权衡。A predictive familiarity estimator that leverages mid-layer hidden states to evaluate semantic competence among agents, avoiding the overhead of full rollouts and achieving a trade-off between task performance and efficiency is designed.
本文提出 Fathom,一种 key scan,其中每个查询决定读取每个 key 通道的比特数;在真实的 coding agent 会话中,Fathom 以 92 bit 达到最准确的 136 bit scan 的步骤一致性。Fathom is presented, a key scan in which each query decides how many bits of each key channel to read, and on real coding-agent sessions Fathom reaches the step agreement of the most accurate 136-bit scan at 92 bits.
BraTS-GoAT 在异质人群上使用传统 3D nnU-Net 对 1,351 个标注案例进行肿瘤分割评估,采用五折交叉验证,每折 1,000 epochs,并应用 test-time mirroring。BraTS-GoAT evaluates tumor segmentation across heterogeneous populations using a conventional 3D nnU-Net on 1,351 labeled cases using five-fold cross-validation and 1,000 epochs per fold and applied test-time mirroring.
常用的并行方案留下了四种未被约束的并行维度,各自增长方式不同:随路由矩阵增长的专家调度、随 token 数乘词表规模增长的词表投影、随深度乘序列长度增长的梯度检查点边界,以及随参数量增长的优化器状态。Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.
该工作提出 RetireOPD(Self-Retiring On-Policy Distillation),先用环境奖励优化一个解耦的、技能条件化的教师模型,再联合 RL 与 OPD 训练一个无技能依赖的学生模型。This work proposes RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.
提出 WeVisDoc,一个面向鲁棒端到端文档解析的两阶段数据驱动框架,通过异构数据与保持结构的退化合成来扩展语义、结构与外观覆盖度,并指导有针对性的数据构建与目标 token 预算的重分配。WeVisDoc is presented, a two-stage data-centric framework for robust end-to-end document parsing that broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis and guides targeted data construction and reallocation of the target-token budget.
构建了 AMPLE-Math,一个包含 5,319 道数学问题、每题对应六种共享同一答案的推理视角的可复用套件,并通过对比表明 OPSD 能借助直答推理与带思维链推理所共享的参数,提升对已有推理能力的调用效率。AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is constructed and AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is compared, suggesting that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference.
提出 UFO,这是首个面向全条件对齐同时评估的统一框架;并发布 UFO-Bench,一个用于整体评估现有定制化模型在文本与视觉条件多样化交互下表现的专用基准。UFO is proposed, the first unified framework for omni-condition alignment simultaneous evaluation, and UFO-Bench is presented, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.
本文提出 When2Think,一种基于 RLVR 的后训练框架,用于实例自适应计算分配,既无需学习奖励模型,也无需学习 critic;离线参考缓存机制避免了策略更新阶段对参考模型的在线查询。This work proposes When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation that requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates.
我们研究 on-policy 蒸馏 (OPD) 中的长度膨胀现象,即学生回答会变得过长,甚至耗尽生成预算。我们发现基础学生模型与训练后教师模型之间的终止 token 不匹配是该行为的重要来源。在 Qwen3、Llama 和 Gemma 上,两个模型可能将停止概率分配到不同的 EOS token 上,即便它们声明的停止集合相同。这种不匹配会抑制学生偏好的终止动作,同时无法可靠地传递教师偏好的替代动作。我们证明对齐 t...We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning t
本文提出 FAMOS,一个前馈模型,可从稀疏、无序的部分点云集合预测可动部件分割与关节参数;并引入一个过程式数据生成器,在训练过程中合成自标注资产,以克服现有数据集规模和多样性的局限。FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds, is presented and a procedural data generator that synthesizes self-annotated assets during training is introduced to overcome the limited scale and diversity of existing datasets.
本文通过大规模超参数搜索,系统地重新评估了 Vanilla LoRA 以及九个代表性 LoRA 变体,发现不同 LoRA 方法偏好的学习率区间各异,并将最优学习率区间的差异归因于最大 Hessian 特征值的变化,与经典学习理论相吻合。This work systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches, finding that different LoRA methods favor distinct learning rate ranges and attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.
本文提出 SELF-INDEX,一个无需人工干预即可让索引自我演进的框架;其 Optimizer 可自主诊断检索短板,选择性修改对应的索引键,并在每次更新前对每项修订进行验证。SELF-INDEX is proposed, a framework that enables an index to self-evolve without human intervention, and its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index.
结果表明,仅靠候选数量不足以刻画多候选 test-time scaling 的系统成本;评测应同时报告候选数量与准确率,以及生成调度和 GPU 层级的系统指标。The results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling and Evaluations should report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.
本文提出 Fuse——一个用于研究用户中介社会推理的多智能体仿真框架,并将其应用于 12 个 LLM,通过系统性隔离关键因素来展示其分析效用。This work introduces Fuse, a multi-agent simulation framework for studying user-mediated social reasoning, and applies Fuse to 12 LLMs and demonstrates its analytical utility by systematically isolating key factors.
提出 Calibrated On-Policy Distillation:通过正、负特权干预估计教师的自偏离区间,并仅保留超出该区间的部分以校准原始的 teacher–student 差异。Calibrated On-Policy Distillation is introduced, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region.
提出 Paint-Anything,通过物体级颜色监督学习一个统一的 hex-prompt 界面以同时支持生成与编辑;并引入 Any Color Benchmark (ACBench),包含 ACBench-T2I 和 ACBench-Edit,以衡量两项任务中的物体级 hex 颜色保真度。Paint-Anything is presented, which learns a shared hex-prompt interface for generation and editing through object-level color supervision, and introduces Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks.
结果表明,DeformSmith 生成的资产在视觉质量与物理合理性上均优于 SOTA 基线,包括 PhysGen3D、PhysGM 和 PhysX-Omni,同时支持为可形变物体的机器人操作合成训练数据。Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects.
本文使用 Fast Iterative Shrinkage-Thresholding Algorithm 展开 CSC 优化,并将稀疏系数视为可微变量与网络参数联合学习;同时提出一种 label-free 的后训练策略,在固定主网络参数的情况下,根据被损坏输入自适应调整压缩强度。This work unfolds the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm and treats the sparsity coefficient as a differentiable variable jointly learned with the network parameters and introduces a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed.
Mixture of Memory Embeddings (MoME):一种上下文感知的记忆机制,将每个 token 的单一记忆行替换为 M 个 slot 的混合,并通过隐藏状态上的可学习门控选择在每个位置读取哪些 slot;在 sub-billion 规模下展现出更优的记忆容量 scaling 趋势,且训练与推理均保持高效。Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position, shows more promising memory-size scaling trend at sub-billion scale and remains efficient in training and inference.
IntBMoE 是一种 block-conditioned MoE,将三者解耦,通过密集专家组合与稀疏 block 执行配对实现;在图像分类任务上,相较代表性的稀疏与密集 MoE 基线均取得稳定提升。IntBMoE is a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution, and shows consistent gains over representative sparse and dense MoE baselines on image classification.
RefineEdit 是一个基于 GRN 的 training-free prompt-to-prompt 图像编辑框架,将 bit routing 与两种稳定机制(adaptive spatial freezing 与 finite bit locking)相结合,使编辑证据能够随图像演化而被修正。RefineEdit is a training-free prompt-to-prompt image editing framework built on the GRN that combines bit routing with two stabilization mechanisms: adaptive spatial freezing and finite bit locking, allowing editing evidence to be revised as the image evolves.
提出一种持续自适应框架:冻结的前沿模型通过超过 230 个工具操作专业设计软件,同时外部的自然语言技能程序性记忆从经验中不断积累并精炼可复用的设计流程。A continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience is introduced.