研究库 论文知识库
Papers · organized/paper_cards

论文

199 张论文卡片 · 工程化

开放获取 全部 绿色 · 1640
条目 G: 公共部门 ML Pipeline 工程教训(含性能数据表)
arXiv:2511.01545 工程化 观点 OA · 绿色 被引 1 · S2

研究表明,机器学习在公共部门的成功,将更少依赖模型准确率的突破,而更多依赖机构能否构建出透明、可复现、可问责且受公民信任的数据基础设施。It is shown that the success of machine learning in the public sector will depend less on breakthroughs in model accuracy and more on the ability of institutions to engineer transparent, reproducible, and accountable data infrastructures that citizens can trust.

条目 F: Google 企业定制 LLM — 代码转换实战数据
arXiv:2605.16517 工程化 方法 OA · 绿色 被引 1 · S2

介绍 Gemini for Google (GfG),一款面向 Google 内部软件工程生态的 Gemini 专用适配版本,涵盖从构建万亿 token 的专有数据集到采用可缓解灾难性遗忘的中段训练策略的完整过程。G Gemini for Google (GfG)}, an adaptation of Gemini specialized for Google's internal software engineering ecosystem, is introduced, from curating a trillion-token proprietary dataset to implementing a mid-training strategy that mitigates catastrophic forgetting.

条目 E-NF2:MLOps 架构指南 — 25 条模型集成/部署规范(灰色文献综述)
arXiv:2606.06535 工程化 综述 OA · 绿色 被引 0 · S2 + OpenAlex

本文贡献了一组在架构层面具有重要意义的 MLOps 集成与部署指南(共 25 条),分为五类,并阐述其对整体系统架构的影响。This work contributes a collection of 25 architecturally significant MLOps guidelines for model integration and deployment, organized into five categories, and describes their impact on the overall system architecture.

⑥ "How are MLOps Frameworks Used in Open Source Projects"(arXiv:2601.18591)
⑥ "How are MLOps Frameworks Used in Open Source Projects"(arXiv:2601.18591)
arXiv:2601.18591 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

对八款主流开源 MLOps 框架的实践使用与功能增强需求进行调查,结果显示 MLOps 框架很少被直接开箱即用,也较少集成进 GitHub Workflows,开发者更多通过其 API 在项目中实现自定义功能。Investigating the practical use and desired feature enhancements of eight popular open-source MLOps frameworks indicates that users mainly ask for enhancements to core features of the frameworks, but also better API exposure and CI/CD integration.

⑤ MLOps系统综述(arXiv:2604.16371)
⑤ MLOps系统综述(arXiv:2604.16371)
arXiv:2604.16371 工程化 综述 OA · 绿色 被引 2 · S2

对聚焦MLOps工具的学术文献进行系统综述,揭示其功能、范围及其旨在解决的挑战,并突出真实MLOps pipeline中各工具间互操作性的重要性。A systematic review of the academic literature focused on MLOps tools is conducted to reveal their function, scope, and the challenges they are designed to address and highlight the importance of interoperability across MLOps tools in real-world MLOps pipelines.

Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification
训练留痕:用于语言模型谱系验证的居中残差签名
arXiv:2608.14929 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

结果为兼容的开源权重语言模型检查点建立了一种被动、无需数据的溯源信号,且该投影配对信号出现在六个及更多语言模型系列中The results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints, and the projection-pairing signal appears across six language-model families and beyond.

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
越流行越难遗忘:面向 LLM 遗忘的自适应流行度方法
arXiv:2608.14229 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 AdaPop(自适应流行度)方法,将局部 token 置信度与源自外部代理的逐事实流行度相关指数相结合,并通过双上升控制器在每个 epoch 调整 retain 惩罚来自动平衡遗忘与保留。The AdaPop (Adaptive Popularity) method is proposed, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy, and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch.

[DataEvolver] Automatic Data Preparation for Large Language Models through Multi-Level Self-Evolving
DataEvolver:基于多层级自演化的 LLM 自动化数据准备
arXiv:2606.07001 工程化 方法 OA · 绿色 被引 3 · S2

实验表明,DataEvolver 显著提升了数据质量,相比在原始数据上训练,下游 LLM 性能平均提升 10%,凸显了 LLM 与数据迭代协同演化的新机遇。Experiments show that DataEvolver substantially improves data quality and achieves an average 10\% gain in downstream LLM performance compared with training on original data, highlighting new opportunities for the iterative co-evolution of LLMs and data.

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
Game2World Engine:解锁真实游戏视频用于世界模型训练
arXiv:2608.24680 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 GameCleaner,一个无需 mask 的游戏 UI 移除模型,结合多模态语义理解与视频编辑能力,整体 VideoReward 较在带 UI 数据上训练的模型提升 6.83%。GameCleaner is proposed, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities and improves overall VideoReward by 6.83% over those trained on UI-overlaid data.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Next-chunk reasoning RL 真的优于 SFT 吗?——在 no-CoT 数据下重新审视训练策略
arXiv:2608.23256 工程化 方法 OA · 绿色 被引 1 · S2

Mixed SFT 是一种单阶段监督微调,联合在 no-CoT 和 long-CoT 数据上训练,相比 next-chunk reasoning RL 取得了明显更高的 RLVR 后性能上限,同时训练算力开销减少超过 60 倍。Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.

TTPO: Test-Time Policy Optimization
TTPO:测试时策略优化。
arXiv:2608.27448 工程化 方法 OA · 绿色 被引 1 · S2

提出 Test-Time Policy Optimization,一种非对称目标,通过 OPSD 蒸馏一致性 rollout,并使用 Grouped RL 惩罚不一致的 rollout;进一步通过 token 级选择精炼两个分支:蒸馏降低已收敛位置的权重,而 RL 仅惩罚置信的错误。Test-Time Policy Optimization is proposed, an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL and Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors.

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
获取、修复、保留:面向小模型对话游戏 Agent 的诊断驱动后训练方案
arXiv:2608.28458 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,广泛的 SFT 带来模型大部分能力提升;当失败检测精准时,turn-local 监督可发挥作用,且观察到的迁移主要集中在同族模型之间。The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.

Verification-Aware Training for Speculative Decoding
面向 Speculative Decoding 的验证感知训练
arXiv:2608.30135 工程化 观点 OA · 绿色 被引 2 · S2

提出 Verification-Aware Training,一种插件式框架,在每一步训练中模拟验证并将产生的 accept 与 reject 模式转化为监督信号,在数学、代码和聊天基准上提升了平均接受长度和实际加速比。Verification-Aware Training is introduced, a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision and improves average acceptance length and wall-clock speedup across math, code, and chat benchmarks.

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
从生产流量到后训练:构建覆盖企业请求组合的自托管 LLM
arXiv:2609.01572 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

数据驻留约束迫使企业自托管 LLM,但不断引入新模型而不下线旧模型会扩张服务集群,分散有限的 GPU 池。我们通过沿指令遵循、函数调用和内部任务分布三个维度,针对生产错误分析所发现的质量差距,将 200 多个内部应用的流量整合到单一模型上。质量通过按生产流量分层的离线基准进行跟踪,并由确定性验证器或经过校准的 LLM 评判器打分。不同于针对Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimi

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
中期训练阶段的知识蒸馏更偏向推理而非事实记忆
arXiv:2609.01532 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Switch Distillation,一种简单的 mid-training 目标:以教师预测熵作为轻量路由信号,仅在教师置信的 token 上蒸馏,其余回退到交叉熵;在不同教师规模下均稳定优于现有蒸馏目标。Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
机构报纸流水线:从历史报纸中提炼数十亿高质量 token
arXiv:2608.18972 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Institutional Newspapers Pipeline,一个模块化系统,旨在从历史报纸扫描件中提取高质量、结构化的数据集;其架构设计使每个步骤都保持可解释和可定制,并使整个 pipeline 在计算上足够精简,可在工作站级硬件上运行。The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
Debias-SparseGPT:面向 LLM 的偏置感知剪枝
arXiv:2609.02496 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Debias-SparseGPT,一种 post-training 剪枝方法,通过在人口统计对比输入上定义的二阶项引入表示去偏好,在保持模型困惑度与零样本准确率的前提下,一致地降低剪枝带来的偏差,效果优于 SparseGPT。Debias-SparseGPT is introduced, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs that consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy.

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Compile by Training:将自然语言规约转化为本地神经函数
arXiv:2609.04199 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

许多重复出现的文本函数易于描述却难以用规则实现;而为每个输入调用大型远程模型会带来重复开销、延迟与对服务方的依赖。我们提出 compile by training,将自然语言规约转化为可复用的神经函数。在编译时,教师模型生成任务专属样本,用于为精简解释器训练一个小型适配器。生成的函数可在没有教师模型的情况下运行,并能像普通软件一样被存储、版本化管理与组合。在 FuzzyBench-Hard 这一子集上……Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
知道何时不复用:自主 LLM 后训练中的条件经验迁移
arXiv:2608.26730 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将条件经验迁移进行形式化,并提出 Boundary-Calibrated Intervention Transfer,一种在权重变化的训练之前即授权经验复用的方法,在相同预算下取得比所评估替代方案更高的最终模型质量。Conditional experience transfer is formulated as conditional experience transfer and Boundary-Calibrated Intervention Transfer is introduced, a method that authorizes experience reuse before weight-changing training and attains higher equal-budget final-model quality than the evaluated alternatives.

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Scal3R:学习高效多相对位姿查询以实现可扩展的在线 3D 重建
arXiv:2609.04201 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该方法将在线重建重构为多参考相对位姿查询,在 Virtual KITTI、Sintel、TUM-Dynamic、ScanNet 和 7-Scenes 上取得 SOTA 性能,并在单 GPU 上 8 小时内收敛。This approach reformulates online reconstruction as multi-reference relative pose querying, which achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes and reaches convergence in 8 hours on a single GPU.

4️⃣ arXiv · LoRAFusion(⭐⭐⭐ 值得追踪)
4️⃣ arXiv · LoRAFusion(⭐⭐⭐ 值得追踪)
arXiv:2510.00206 工程化 方法 被引 7 · S2

本文提出 LoRAFusion,一种面向 LLM 的高效 LoRA 微调系统,可消除不必要的内存访问,在不付出重算或同步代价的前提下保持 compute-bound GEMM 的性能,并引入面向多任务微调的自适应批处理算法。LoRAFusion is introduced, an efficient LoRA fine-tuning system for LLMs that eliminates unnecessary memory accesses and preserves the performance of compute-bound GEMMs without incurring the cost of recomputation or synchronization and introduces an adaptive batching algorithm for multi-job fine-tuning.

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
RISE:通过自外推策略蒸馏实现递归改进
arXiv:2609.05295 工程化 方法 OA · 绿色 被引 2 · S2

涵盖数学推理、多领域 STEM、代码生成以及多轮 Agent 任务的实验表明,RISE 在所有设置下均优于仅使用 RLVR 的训练以及 on-policy self-distillation。Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.

Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
面向遥感变化检测的基于真实世界知识引导的变化数据合成
arXiv:2608.24263 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 KnowChange,一个知识引导的变化数据合成框架,利用预训练视觉语言模型作为知识源,从变化前场景与期望变化类型推理合理的变化位置与类别转移,在统一框架下灵活合成多样变化类型。KnowChange is introduced, a knowledge-guided change data synthesis framework that leverages pretrained vision-language models as knowledge sources to reason about plausible change locations and class transitions from pre-change scenes and desired change types and enables flexible synthesis of diverse change types within a unified framework.

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
Split-LLM 训练中的隐私失败:返回的梯度使诱饵失效
arXiv:2609.04382 工程化 方法 OA · 绿色 被引 1 · S2

本文对一个双节点 split-LLM 训练系统进行系统安全案例研究:其隐私评估通过,却遗留一条未被测试的可观测信道;系统因此并不安全——包括跨训练步骤累积观测在内的五类攻击从未被测量。A systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested, but the system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured.

3️⃣ arXiv · MatryoshkaLoRA(⭐⭐⭐⭐ 值得关注)
3️⃣ arXiv · MatryoshkaLoRA(⭐⭐⭐⭐ 值得关注)
arXiv:2605.07850 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MatryoshkaLoRA,一种受 Matryoshka 启发、面向 LoRA 的通用训练框架,通过在已有 LoRA adapter 之间插入一个固定的、经精心设计的对角矩阵来按比例缩放其子秩,从而学习到准确的层次化低秩表示。MatryoshkaLoRA is proposed, a general, Matryoshka-inspired training framework for LoRA that learns accurate hierarchical low-rank representations by inserting a fixed, carefully crafted diagonal matrix between the existing LoRA adapters to scale their sub-ranks accordingly.

Causal Foundation Models
因果基础模型
arXiv:2609.03003 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

因果基础模型是经过预训练的神经网络,可在全新的数据集上通过上下文学习估计因果量(如平均处理效应),无需模型更新。Causal foundation models are pretrained neural networks that estimate causal quantities, such as the average treatment effect, on entirely new datasets using in-context learning without requiring model updates.

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
CoVeR:基于覆盖度的 token 剪枝,用于 VLM 中的多视图 3D 推理
arXiv:2609.08345 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文证明空间覆盖与三维推理性能相关,并提出 CoVeR,一种仅使用 token 坐标、不依赖学习信号的确定性、无训练选择器,在三项三维推理基准上均超越既有 SOTA,并可作为即插即用模块泛化到四种 VLM。It is shown that spatial coverage is associated with 3D reasoning performance and CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals is introduced, which outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs.

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
OpenWAM:面向系统化世界-动作模型预训练的开源模块化探索
arXiv:2609.07398 工程化 方法 OA · 绿色 被引 14 · S2

提出 OpenWAM,一个开源研究栈,将世界-动作预训练转化为可控的实验项目,并发布完整栈,包括基础设施、评估协议、预训练模型和数据配方,以促进未来研究。This work introduces OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program, and releases the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
RenderFormer-V2:基于异构场景基元的神经渲染
arXiv:2609.05738 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 RenderFormer-V2,一个统一的基于 transformer 的学习型神经渲染模型,可与现代基于物理的渲染系统互补,无需逐场景训练或专用代码,即可处理焦散、体积散射、环境光照、带纹理与置换的表面以及分布外材质等多种光传输效果。RenderFormer-V2 将全局光传输建模为序列到序列变换。沿袭前作,它仍采用两阶段流程:先是与视图无关的阶段,解析场景内基元到……We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to

Graph Machine: Towards Better Pretraining via Edges
Graph Machine:通过边实现更优的预训练
arXiv:2609.02881 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Graph Machine,一种保持 O(n) 规模状态并通过稀疏动态路由访问的架构,使用边——由类似指针追逐的引用机制以可微分方式更新的指针类对象。The Graph Machine is introduced, an architecture that maintains an O(n)-sized state and accesses it through sparse, dynamic routing and uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing.

3. Kubernetes for GenAI Inference(arXiv:2602.04900v2)
Kubernetes for GenAI Inference(arXiv:2602.04900v2)
arXiv:2602.04900 工程化 方法 Open MIND OA · 绿色 被引 2 · S2

这些结果表明,Kueue、DAS 与 GAIE 等互补组件构成了一个高性能的协同平台,证明了 Kubernetes 能够作为承载高要求 GenAI 工作负载的统一底座。These findings illustrate that these complementary components (Kueue, DAS, and GAIE) form a cohesive, high-performance platform, proving Kubernetes' capability to serve as a unified foundation for demanding GenAI workloads.

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
从重加权到改写:解锁训练数据归因中有影响力样本的干预效果
arXiv:2609.02771 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文引入影响引导的响应改写方法,利用 IF 识别干预目标,在保持指令不变的情况下将其响应替换为行为对齐或行为对立的监督信号,以此推动对 TDA 方法的干预感知评估。Influence-guided response rewriting is introduced, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed, motivating intervention-aware evaluation of TDA methods.

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
通往 IMO 金牌的开放配方:面向奥林匹克数学的 Nemotron 训练
arXiv:2609.10712 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一个用于困难奥林匹克数学自然语言证明生成的开放模型测试时计算流水线,完全在自然语言中运行,无需形式化证明器、外部工具或互联网访问。An open-model test-time-compute pipeline for natural-language proof generation for hard olympiad mathematics that operates entirely in natural language, with no formal prover, external tools, or internet access is presented.

UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
UniH^3:面向全能医学图像复原的层次同质性与异质性统一
arXiv:2609.11156 工程化 方法 被引 0 · S2

提出 UniH3,一个统一层次同质性与异质性的全新框架,用于一体化医学图像修复,在两个大规模基准上的大量实验表明 UniH3 在一体化和单任务医学图像修复上均达到 SOTA 性能。UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity and Heterogeneity for all-in-one medical image restoration, is proposed and extensive experiments on two large-scale benchmarks demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration.

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
打破视觉-动作捷径:面向可泛化机器人基础模型的潜在接口训练
arXiv:2609.12641 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

LIT(Latent Interface Training)是一个与框架无关的两阶段策略:先在无图像条件下建立空间目标条件化的动作先验,再通过姿态监督的潜在接口约束视觉条件化,可在保持或提升 LIBERO 平均成功率的同时改善 LIBERO-Plus 综合成功率。Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
SAS:通过端到端上下文排序优化实现简单注意力稀疏化
arXiv:2609.13141 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在推理、长上下文理解和 Agent 任务中,SAS 在不同注意力预算下均优于可训练的稀疏注意力基线,在紧预算下增益尤为显著,表明其上下文排序对下游任务更有效。Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.