研究库 论文知识库
Papers · organized/paper_cards

论文

9 张论文卡片 · 工程化 · 观点 · OA 绿色

开放获取 全部 绿色 · 1640
条目 G: 公共部门 ML Pipeline 工程教训(含性能数据表)
arXiv:2511.01545 工程化 观点 OA · 绿色 被引 1 · S2

研究表明,机器学习在公共部门的成功,将更少依赖模型准确率的突破,而更多依赖机构能否构建出透明、可复现、可问责且受公民信任的数据基础设施。It is shown that the success of machine learning in the public sector will depend less on breakthroughs in model accuracy and more on the ability of institutions to engineer transparent, reproducible, and accountable data infrastructures that citizens can trust.

Verification-Aware Training for Speculative Decoding
面向 Speculative Decoding 的验证感知训练
arXiv:2608.30135 工程化 观点 OA · 绿色 被引 2 · S2

提出 Verification-Aware Training,一种插件式框架,在每一步训练中模拟验证并将产生的 accept 与 reject 模式转化为监督信号,在数学、代码和聊天基准上提升了平均接受长度和实际加速比。Verification-Aware Training is introduced, a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision and improves average acceptance length and wall-clock speedup across math, code, and chat benchmarks.

Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand
手指作为腿:使用仿人手学习自支撑运动与操作
arXiv:2609.17172 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该工作展示了一个紧凑的移动机械手,复用其手指同时完成运动与交互,无需独立的运动机构。This work demonstrates a compact mobile manipulator that reuses its fingers for locomotion and interaction, without a separate locomotion mechanism.

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
你的 Transformer 可同时容纳两个思考:LLM 中线性叠加的证据
arXiv:2609.29845 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文给出证据,说明 superposition 是 Transformer 架构的内在属性,而非训练过程中涌现的结果,并证明通过轻量级 fine-tuning 可以在很大程度上恢复线性性。This work provides evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training, and demonstrates that linearity can be substantially restored through lightweight fine-tuning.

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
[标题中文] MOPD-Router:重新思考多教师在策略蒸馏中的教师路由
arXiv:2609.30837 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ExpertAlign,一种无需域标签或独立 routing 模型的框架,可在每个 token 上对全体 teacher 池进行监督路由;研究表明 token 级路由能够利用跨域互补监督,并减少对 prompt 级域指派的单一依赖。This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment.

Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
通过位置选择性自蒸馏从语言反馈训练 LLM 评判器
arXiv:2609.38792 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,屏蔽高熵偏移位置相比朴素的 SD 提升了分布外泛化能力,由此得到的自蒸馏 judges 在所评估的主观子类别上比基于结果监督 RL 训练的 judges 高出 2-9 个百分点,同时在客观子类别上保持竞争力。Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD, and the resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.

WARP: Weight-Space Analysis for Recovering Training Data Portfolios
WARP:基于权重空间分析的训练数据组合还原
arXiv:2607.01686 工程化 观点 OA · 绿色 被引 1 · S2

WARP:一个直接从已发布权重还原微调模型训练数据混合比例的框架,抽取几何特征并映射至各领域占比,可采用无参数 softmax 读出器,或基于合成混合训练的 MLP 投影器。WARP is introduced, a framework that recovers a fine-tuned model's training mixtures directly from its released weights and extracts geometric features and maps them to domain proportions using either a parameter-free softmax readout or an MLP projector trained on synthetic mixtures.

Randomized YaRN Improves Length Generalization for Long-Context Reasoning
Randomized YaRN 改善长上下文推理的长度泛化能力
arXiv:2606.23687 工程化 观点 OA · 绿色 被引 1 · S2

本文提出 Randomized YaRN,一种通过将基于 YaRN 的位置外推与随机位置编码和长度课程相结合来提升长度泛化能力的训练方法,表明渐进式地将模型暴露于分布外位置分布是实现可泛化长上下文推理的有效方案。Randomized YaRN is proposed, a training method that improves length generalization by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum, and suggests that progressively exposing models to OOD positional distributions provides an effective recipe for generalizable long-context reasoning.

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
当注意力"失明":ALiBi 位置编码中的数值失效
arXiv:2608.03994 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

发现 ALiBi 的失效模式会显著损害 token 检索,而对标准 decoder 基准影响较小;提出四种训练时缓解策略,在 passkey 检索上获得最一致的提升。It is found that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks, and proposes four training-time mitigation strategies that yield the most consistent improvements in passkey retrieval.