研究库 论文知识库
Papers · organized/paper_cards

论文

148 张论文卡片 · 工程化 · 方法

开放获取 全部 绿色 · 1640
条目 F: Google 企业定制 LLM — 代码转换实战数据
arXiv:2605.16517 工程化 方法 OA · 绿色 被引 1 · S2

介绍 Gemini for Google (GfG),一款面向 Google 内部软件工程生态的 Gemini 专用适配版本,涵盖从构建万亿 token 的专有数据集到采用可缓解灾难性遗忘的中段训练策略的完整过程。G Gemini for Google (GfG)}, an adaptation of Gemini specialized for Google's internal software engineering ecosystem, is introduced, from curating a trillion-token proprietary dataset to implementing a mid-training strategy that mitigates catastrophic forgetting.

⑥ "How are MLOps Frameworks Used in Open Source Projects"(arXiv:2601.18591)
⑥ "How are MLOps Frameworks Used in Open Source Projects"(arXiv:2601.18591)
arXiv:2601.18591 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

对八款主流开源 MLOps 框架的实践使用与功能增强需求进行调查,结果显示 MLOps 框架很少被直接开箱即用,也较少集成进 GitHub Workflows,开发者更多通过其 API 在项目中实现自定义功能。Investigating the practical use and desired feature enhancements of eight popular open-source MLOps frameworks indicates that users mainly ask for enhancements to core features of the frameworks, but also better API exposure and CI/CD integration.

[DataEvolver] Automatic Data Preparation for Large Language Models through Multi-Level Self-Evolving
DataEvolver:基于多层级自演化的 LLM 自动化数据准备
arXiv:2606.07001 工程化 方法 OA · 绿色 被引 3 · S2

实验表明,DataEvolver 显著提升了数据质量,相比在原始数据上训练,下游 LLM 性能平均提升 10%,凸显了 LLM 与数据迭代协同演化的新机遇。Experiments show that DataEvolver substantially improves data quality and achieves an average 10\% gain in downstream LLM performance compared with training on original data, highlighting new opportunities for the iterative co-evolution of LLMs and data.

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
Game2World Engine:解锁真实游戏视频用于世界模型训练
arXiv:2608.24680 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 GameCleaner,一个无需 mask 的游戏 UI 移除模型,结合多模态语义理解与视频编辑能力,整体 VideoReward 较在带 UI 数据上训练的模型提升 6.83%。GameCleaner is proposed, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities and improves overall VideoReward by 6.83% over those trained on UI-overlaid data.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Next-chunk reasoning RL 真的优于 SFT 吗?——在 no-CoT 数据下重新审视训练策略
arXiv:2608.23256 工程化 方法 OA · 绿色 被引 1 · S2

Mixed SFT 是一种单阶段监督微调,联合在 no-CoT 和 long-CoT 数据上训练,相比 next-chunk reasoning RL 取得了明显更高的 RLVR 后性能上限,同时训练算力开销减少超过 60 倍。Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.

TTPO: Test-Time Policy Optimization
TTPO:测试时策略优化。
arXiv:2608.27448 工程化 方法 OA · 绿色 被引 1 · S2

提出 Test-Time Policy Optimization,一种非对称目标,通过 OPSD 蒸馏一致性 rollout,并使用 Grouped RL 惩罚不一致的 rollout;进一步通过 token 级选择精炼两个分支:蒸馏降低已收敛位置的权重,而 RL 仅惩罚置信的错误。Test-Time Policy Optimization is proposed, an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL and Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors.

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
中期训练阶段的知识蒸馏更偏向推理而非事实记忆
arXiv:2609.01532 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Switch Distillation,一种简单的 mid-training 目标:以教师预测熵作为轻量路由信号,仅在教师置信的 token 上蒸馏,其余回退到交叉熵;在不同教师规模下均稳定优于现有蒸馏目标。Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
机构报纸流水线:从历史报纸中提炼数十亿高质量 token
arXiv:2608.18972 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Institutional Newspapers Pipeline,一个模块化系统,旨在从历史报纸扫描件中提取高质量、结构化的数据集;其架构设计使每个步骤都保持可解释和可定制,并使整个 pipeline 在计算上足够精简,可在工作站级硬件上运行。The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Compile by Training:将自然语言规约转化为本地神经函数
arXiv:2609.04199 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

许多重复出现的文本函数易于描述却难以用规则实现;而为每个输入调用大型远程模型会带来重复开销、延迟与对服务方的依赖。我们提出 compile by training,将自然语言规约转化为可复用的神经函数。在编译时,教师模型生成任务专属样本,用于为精简解释器训练一个小型适配器。生成的函数可在没有教师模型的情况下运行,并能像普通软件一样被存储、版本化管理与组合。在 FuzzyBench-Hard 这一子集上……Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
知道何时不复用:自主 LLM 后训练中的条件经验迁移
arXiv:2608.26730 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将条件经验迁移进行形式化,并提出 Boundary-Calibrated Intervention Transfer,一种在权重变化的训练之前即授权经验复用的方法,在相同预算下取得比所评估替代方案更高的最终模型质量。Conditional experience transfer is formulated as conditional experience transfer and Boundary-Calibrated Intervention Transfer is introduced, a method that authorizes experience reuse before weight-changing training and attains higher equal-budget final-model quality than the evaluated alternatives.

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Scal3R:学习高效多相对位姿查询以实现可扩展的在线 3D 重建
arXiv:2609.04201 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该方法将在线重建重构为多参考相对位姿查询,在 Virtual KITTI、Sintel、TUM-Dynamic、ScanNet 和 7-Scenes 上取得 SOTA 性能,并在单 GPU 上 8 小时内收敛。This approach reformulates online reconstruction as multi-reference relative pose querying, which achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes and reaches convergence in 8 hours on a single GPU.

4️⃣ arXiv · LoRAFusion(⭐⭐⭐ 值得追踪)
4️⃣ arXiv · LoRAFusion(⭐⭐⭐ 值得追踪)
arXiv:2510.00206 工程化 方法 被引 7 · S2

本文提出 LoRAFusion,一种面向 LLM 的高效 LoRA 微调系统,可消除不必要的内存访问,在不付出重算或同步代价的前提下保持 compute-bound GEMM 的性能,并引入面向多任务微调的自适应批处理算法。LoRAFusion is introduced, an efficient LoRA fine-tuning system for LLMs that eliminates unnecessary memory accesses and preserves the performance of compute-bound GEMMs without incurring the cost of recomputation or synchronization and introduces an adaptive batching algorithm for multi-job fine-tuning.

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
RISE:通过自外推策略蒸馏实现递归改进
arXiv:2609.05295 工程化 方法 OA · 绿色 被引 2 · S2

涵盖数学推理、多领域 STEM、代码生成以及多轮 Agent 任务的实验表明,RISE 在所有设置下均优于仅使用 RLVR 的训练以及 on-policy self-distillation。Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.

Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
面向遥感变化检测的基于真实世界知识引导的变化数据合成
arXiv:2608.24263 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 KnowChange,一个知识引导的变化数据合成框架,利用预训练视觉语言模型作为知识源,从变化前场景与期望变化类型推理合理的变化位置与类别转移,在统一框架下灵活合成多样变化类型。KnowChange is introduced, a knowledge-guided change data synthesis framework that leverages pretrained vision-language models as knowledge sources to reason about plausible change locations and class transitions from pre-change scenes and desired change types and enables flexible synthesis of diverse change types within a unified framework.

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
Split-LLM 训练中的隐私失败:返回的梯度使诱饵失效
arXiv:2609.04382 工程化 方法 OA · 绿色 被引 1 · S2

本文对一个双节点 split-LLM 训练系统进行系统安全案例研究:其隐私评估通过,却遗留一条未被测试的可观测信道;系统因此并不安全——包括跨训练步骤累积观测在内的五类攻击从未被测量。A systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested, but the system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured.

3️⃣ arXiv · MatryoshkaLoRA(⭐⭐⭐⭐ 值得关注)
3️⃣ arXiv · MatryoshkaLoRA(⭐⭐⭐⭐ 值得关注)
arXiv:2605.07850 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MatryoshkaLoRA,一种受 Matryoshka 启发、面向 LoRA 的通用训练框架,通过在已有 LoRA adapter 之间插入一个固定的、经精心设计的对角矩阵来按比例缩放其子秩,从而学习到准确的层次化低秩表示。MatryoshkaLoRA is proposed, a general, Matryoshka-inspired training framework for LoRA that learns accurate hierarchical low-rank representations by inserting a fixed, carefully crafted diagonal matrix between the existing LoRA adapters to scale their sub-ranks accordingly.

Causal Foundation Models
因果基础模型
arXiv:2609.03003 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

因果基础模型是经过预训练的神经网络,可在全新的数据集上通过上下文学习估计因果量(如平均处理效应),无需模型更新。Causal foundation models are pretrained neural networks that estimate causal quantities, such as the average treatment effect, on entirely new datasets using in-context learning without requiring model updates.

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
CoVeR:基于覆盖度的 token 剪枝,用于 VLM 中的多视图 3D 推理
arXiv:2609.08345 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文证明空间覆盖与三维推理性能相关,并提出 CoVeR,一种仅使用 token 坐标、不依赖学习信号的确定性、无训练选择器,在三项三维推理基准上均超越既有 SOTA,并可作为即插即用模块泛化到四种 VLM。It is shown that spatial coverage is associated with 3D reasoning performance and CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals is introduced, which outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs.

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
OpenWAM:面向系统化世界-动作模型预训练的开源模块化探索
arXiv:2609.07398 工程化 方法 OA · 绿色 被引 14 · S2

提出 OpenWAM,一个开源研究栈,将世界-动作预训练转化为可控的实验项目,并发布完整栈,包括基础设施、评估协议、预训练模型和数据配方,以促进未来研究。This work introduces OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program, and releases the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
RenderFormer-V2:基于异构场景基元的神经渲染
arXiv:2609.05738 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 RenderFormer-V2,一个统一的基于 transformer 的学习型神经渲染模型,可与现代基于物理的渲染系统互补,无需逐场景训练或专用代码,即可处理焦散、体积散射、环境光照、带纹理与置换的表面以及分布外材质等多种光传输效果。RenderFormer-V2 将全局光传输建模为序列到序列变换。沿袭前作,它仍采用两阶段流程:先是与视图无关的阶段,解析场景内基元到……We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to

Graph Machine: Towards Better Pretraining via Edges
Graph Machine:通过边实现更优的预训练
arXiv:2609.02881 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Graph Machine,一种保持 O(n) 规模状态并通过稀疏动态路由访问的架构,使用边——由类似指针追逐的引用机制以可微分方式更新的指针类对象。The Graph Machine is introduced, an architecture that maintains an O(n)-sized state and accesses it through sparse, dynamic routing and uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing.

3. Kubernetes for GenAI Inference(arXiv:2602.04900v2)
Kubernetes for GenAI Inference(arXiv:2602.04900v2)
arXiv:2602.04900 工程化 方法 Open MIND OA · 绿色 被引 2 · S2

这些结果表明,Kueue、DAS 与 GAIE 等互补组件构成了一个高性能的协同平台,证明了 Kubernetes 能够作为承载高要求 GenAI 工作负载的统一底座。These findings illustrate that these complementary components (Kueue, DAS, and GAIE) form a cohesive, high-performance platform, proving Kubernetes' capability to serve as a unified foundation for demanding GenAI workloads.

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
从重加权到改写:解锁训练数据归因中有影响力样本的干预效果
arXiv:2609.02771 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文引入影响引导的响应改写方法,利用 IF 识别干预目标,在保持指令不变的情况下将其响应替换为行为对齐或行为对立的监督信号,以此推动对 TDA 方法的干预感知评估。Influence-guided response rewriting is introduced, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed, motivating intervention-aware evaluation of TDA methods.

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
通往 IMO 金牌的开放配方:面向奥林匹克数学的 Nemotron 训练
arXiv:2609.10712 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一个用于困难奥林匹克数学自然语言证明生成的开放模型测试时计算流水线,完全在自然语言中运行,无需形式化证明器、外部工具或互联网访问。An open-model test-time-compute pipeline for natural-language proof generation for hard olympiad mathematics that operates entirely in natural language, with no formal prover, external tools, or internet access is presented.

UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
UniH^3:面向全能医学图像复原的层次同质性与异质性统一
arXiv:2609.11156 工程化 方法 被引 0 · S2

提出 UniH3,一个统一层次同质性与异质性的全新框架,用于一体化医学图像修复,在两个大规模基准上的大量实验表明 UniH3 在一体化和单任务医学图像修复上均达到 SOTA 性能。UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity and Heterogeneity for all-in-one medical image restoration, is proposed and extensive experiments on two large-scale benchmarks demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration.

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
打破视觉-动作捷径:面向可泛化机器人基础模型的潜在接口训练
arXiv:2609.12641 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

LIT(Latent Interface Training)是一个与框架无关的两阶段策略:先在无图像条件下建立空间目标条件化的动作先验,再通过姿态监督的潜在接口约束视觉条件化,可在保持或提升 LIBERO 平均成功率的同时改善 LIBERO-Plus 综合成功率。Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
SAS:通过端到端上下文排序优化实现简单注意力稀疏化
arXiv:2609.13141 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在推理、长上下文理解和 Agent 任务中,SAS 在不同注意力预算下均优于可训练的稀疏注意力基线,在紧预算下增益尤为显著,表明其上下文排序对下游任务更有效。Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
PLC-DPO:噪声与模糊偏好优化中的后验标签修正
arXiv:2608.30597 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

PLC-DPO 将每个偏好对的训练信号路由为 clean、flip 或 tie 三类,将噪声偏好学习从单纯过滤可疑样本重构为主动修正监督方向与强度。PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.

Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Feyospace-v1:Cyber Mercury Seven 如何训练前沿网络安全模型
arXiv:2609.08418 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

这是首个端到端证明:一个七人独立团队可以训练出在 agentic 网络安全能力上领先的开源权重模型,且全部三个 checkpoint 在相近参数规模下均排名第一。This is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability and all three checkpoints rank 1st among models at comparable parameter scales.

Expert-Space Exploration in MoE Reinforcement Learning
MoE 强化学习中的专家空间探索
arXiv:2609.13058 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ESRL,一种架构感知的框架,显式探索 MoE 模型的专家路由空间,将高置信度专家保留为锚点,并把随机路由限制在合理候选池内,从而保留可靠的计算路径。ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.

MInTRL: Off-policy Intervention can boost On-policy RL
MInTRL:Off-policy 干预可增强 On-policy RL
arXiv:2609.12419 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了最小干预强化学习(Minimal Intervention Reinforcement Learning, MInTRL),通过在原本的 on-policy rollout 中引入稀疏的局部干预来扩展探索边界,确立了最小干预作为增强 on-policy RL 的有效范式。This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.

Modality-Autoregressive World-Action Models
模态自回归世界-动作模型
arXiv:2609.17524 工程化 方法 OA · 绿色 被引 1 · S2

论文提出了 ModAR,这是首个在预测动作前以自回归方式对多种未来模态进行去噪的 WAM;研究发现 ModAR 的序列化生成优于现有 WAM 形式,并在所有评估数据规模下取得最高的平均成功率。ModAR is introduced, the first WAM to autoregressively denoise multiple future modalities before predicting actions, and it is found that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales.

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
Mind2Dialogue:通过模拟用户心理状态训练具备人类感知能力的语言模型
arXiv:2609.15972 工程化 方法 OA · 绿色 被引 1 · S2

Mind2Dialogue 框架提出了一个心理学引导的模拟器,在交互过程中保留个人特征并更新心智状态以生成连贯对话;通过对 Oracle 信息充分的回复进行训练,使模型在部署时无需直接访问用户心智状态即可提供帮助。The Mind2Dialogue framework proposes a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations, and trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment.

Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
无推理轨迹的专家模型训练:面向领域专家蒸馏
arXiv:2609.13770 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作表明,专用优化隐式地从此潜空间轨迹中进行选择,并建立了一种新的专用训练视角:当缺少 gold reasoning 时,调参选择直接控制传递给下游模型的潜在监督信号。This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.

LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
LimiX-2:迈向通用结构化数据智能的上下文机制网络
arXiv:2609.17488 工程化 方法 OA · 绿色 被引 4 · S2

我们推出 LimiX 家族新模型 LimiX-2,通过先前建立的 scaling laws 指导模型与数据规模扩展。LimiX-2 采用上下文机制网络 (CMNs) 范式,并以上下文条件掩码建模 (CCMM) 进行预训练。CMNs 将上下文学习的组织原则从以目标为中心的预测转向以机制为导向的联合建模。它并非围绕传统表格 PFN 的 p(y|x, D_context) 目标设计网络,而是围绕学习 p(x, y|D_context)——一种上下文依赖的表征We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the p(y mid x, D_{context}) objective of conventional tabular PFNs, it is designed around learning p(x, y mid D_{context}), a context-dependent representat

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
长上下文 Mixture-of-Experts 训练中每个内存峰值的平整化
arXiv:2609.14306 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

常用的并行方案留下了四种未被约束的并行维度,各自增长方式不同:随路由矩阵增长的专家调度、随 token 数乘词表规模增长的词表投影、随深度乘序列长度增长的梯度检查点边界,以及随参数量增长的优化器状态。Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.