研究库 论文知识库
Papers · organized/paper_cards

论文

188 张论文卡片 · 工程化 · OA 绿色

开放获取 全部 绿色 · 1640
Expert-Space Exploration in MoE Reinforcement Learning
MoE 强化学习中的专家空间探索
arXiv:2609.13058 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ESRL,一种架构感知的框架,显式探索 MoE 模型的专家路由空间,将高置信度专家保留为锚点,并把随机路由限制在合理候选池内,从而保留可靠的计算路径。ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.

MInTRL: Off-policy Intervention can boost On-policy RL
MInTRL:Off-policy 干预可增强 On-policy RL
arXiv:2609.12419 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了最小干预强化学习(Minimal Intervention Reinforcement Learning, MInTRL),通过在原本的 on-policy rollout 中引入稀疏的局部干预来扩展探索边界,确立了最小干预作为增强 on-policy RL 的有效范式。This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.

Modality-Autoregressive World-Action Models
模态自回归世界-动作模型
arXiv:2609.17524 工程化 方法 OA · 绿色 被引 1 · S2

论文提出了 ModAR,这是首个在预测动作前以自回归方式对多种未来模态进行去噪的 WAM;研究发现 ModAR 的序列化生成优于现有 WAM 形式,并在所有评估数据规模下取得最高的平均成功率。ModAR is introduced, the first WAM to autoregressively denoise multiple future modalities before predicting actions, and it is found that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales.

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
Mind2Dialogue:通过模拟用户心理状态训练具备人类感知能力的语言模型
arXiv:2609.15972 工程化 方法 OA · 绿色 被引 1 · S2

Mind2Dialogue 框架提出了一个心理学引导的模拟器,在交互过程中保留个人特征并更新心智状态以生成连贯对话;通过对 Oracle 信息充分的回复进行训练,使模型在部署时无需直接访问用户心智状态即可提供帮助。The Mind2Dialogue framework proposes a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations, and trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment.

Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
无推理轨迹的专家模型训练:面向领域专家蒸馏
arXiv:2609.13770 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作表明,专用优化隐式地从此潜空间轨迹中进行选择,并建立了一种新的专用训练视角:当缺少 gold reasoning 时,调参选择直接控制传递给下游模型的潜在监督信号。This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.

LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
LimiX-2:迈向通用结构化数据智能的上下文机制网络
arXiv:2609.17488 工程化 方法 OA · 绿色 被引 4 · S2

我们推出 LimiX 家族新模型 LimiX-2,通过先前建立的 scaling laws 指导模型与数据规模扩展。LimiX-2 采用上下文机制网络 (CMNs) 范式,并以上下文条件掩码建模 (CCMM) 进行预训练。CMNs 将上下文学习的组织原则从以目标为中心的预测转向以机制为导向的联合建模。它并非围绕传统表格 PFN 的 p(y|x, D_context) 目标设计网络,而是围绕学习 p(x, y|D_context)——一种上下文依赖的表征We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the p(y mid x, D_{context}) objective of conventional tabular PFNs, it is designed around learning p(x, y mid D_{context}), a context-dependent representat

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
长上下文 Mixture-of-Experts 训练中每个内存峰值的平整化
arXiv:2609.14306 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

常用的并行方案留下了四种未被约束的并行维度,各自增长方式不同:随路由矩阵增长的专家调度、随 token 数乘词表规模增长的词表投影、随深度乘序列长度增长的梯度检查点边界,以及随参数量增长的优化器状态。Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.

Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand
手指作为腿:使用仿人手学习自支撑运动与操作
arXiv:2609.17172 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该工作展示了一个紧凑的移动机械手,复用其手指同时完成运动与交互,无需独立的运动机构。This work demonstrates a compact mobile manipulator that reuses its fingers for locomotion and interaction, without a separate locomotion mechanism.

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
When2Think:面向高效混合推理模型的难度感知长度控制学习
arXiv:2609.19671 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 When2Think,一种基于 RLVR 的后训练框架,用于实例自适应计算分配,既无需学习奖励模型,也无需学习 critic;离线参考缓存机制避免了策略更新阶段对参考模型的在线查询。This work proposes When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation that requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates.

Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
Srijika:面向九种印度文字的 OpenType 布局复用字体再设计
arXiv:2609.05661 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Srijika,一个为九种 Brahmic 文字(Devanagari、Tamil、Bengali、Telugu、Kannada、Malayalam、Gujarati、Gurmukhi、Odia)生成可安装 OpenType 字体的系统,并附带一份负面结果目录,覆盖参考引导重风格化中失败的 conditioning、目标函数选择和数据凸包限制。Srijika is presented, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia, and a negative-results catalogue covering failed conditioning, objective choices, and data-hull limits of reference-guided restyling.

Calibrating Teacher--Student Discrepancy for On-Policy Distillation
用于 On-Policy Distillation 的教师-学生差异校准
arXiv:2609.21619 工程化 方法 OA · 绿色 被引 2 · S2

提出 Calibrated On-Policy Distillation:通过正、负特权干预估计教师的自偏离区间,并仅保留超出该区间的部分以校准原始的 teacher–student 差异。Calibrated On-Policy Distillation is introduced, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region.

Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
基于信息瓶颈的训练自适应卷积稀疏编码,用于鲁棒视觉表征
arXiv:2609.19122 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文使用 Fast Iterative Shrinkage-Thresholding Algorithm 展开 CSC 优化,并将稀疏系数视为可微变量与网络参数联合学习;同时提出一种 label-free 的后训练策略,在固定主网络参数的情况下,根据被损坏输入自适应调整压缩强度。This work unfolds the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm and treats the sparsity coefficient as a differentiable variable jointly learned with the network parameters and introduces a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed.

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
精化本身即可编辑:基于生成式精化网络的无训练 Prompt-to-Prompt 图像编辑
arXiv:2609.20633 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

RefineEdit 是一个基于 GRN 的 training-free prompt-to-prompt 图像编辑框架,将 bit routing 与两种稳定机制(adaptive spatial freezing 与 finite bit locking)相结合,使编辑证据能够随图像演化而被修正。RefineEdit is a training-free prompt-to-prompt image editing framework built on the GRN that combines bit routing with two stabilization mechanisms: adaptive spatial freezing and finite bit locking, allowing editing evidence to be revised as the image evolves.

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
从预训练到精通:面向长时程操作的真实世界子任务强化学习,最小化人工介入
arXiv:2609.21788 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PARTS(Policy Adaptation with RL on Targeted Subtasks),一种真实世界子任务强化学习框架,能够将训练集中于瓶颈环节,同时以最小的人工干预推进训练 rollout。PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at bottlenecks while allowing training rollouts to proceed with minimal human intervention, is presented.

OmniEdu: Open Foundation Models for Learning and Teaching
OmniEdu: Open Foundation Models for Learning and Teaching
arXiv:2609.23088 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

证明了对通用语言模型适配教育任务(涵盖问题求解、课程理解与教学辅助)的精选、能力均衡的监督数据的价值。The value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support is demonstrated.

ACLArena: Agent Continue Learning in Multi-stage Post-training
ACLArena:多阶段后训练中的 Agent 持续学习
arXiv:2609.23989 工程化 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ACLArena,一个全面研究、分析与评估 Agent 持续学习(ACL)的框架,并提出新的 ACL 方案:将高质量轨迹的离线回放与多个由 RL 专精化的 LoRA 专家路由网络相结合,显著提升 Agent 跨多领域学习的能力。This work introduces ACLArena, a framework for comprehensively studying, analyzing, and evaluating Agent Continual Learning, and proposes a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains.

Towards Full Pipeline FP8 Reinforcement Learning for LLMs
面向 LLM 的全流水线 FP8 强化学习
arXiv:2609.22870 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Calibrated Clipping,一种动态方法,通过匹配下界裁剪分位数并相应重平衡上界,将 FP8 裁剪边界与高精度 BF16 分布对齐,消除熵激增并恢复与 BF16 基线可比的性能。Calibrated Clipping is proposed, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly, which eliminates entropy surges and restores performance comparable to the BF16 baseline.

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
StableVQ:稳定的 Vector-Quantized Tokenizer 训练实用指南
arXiv:2609.26774 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 StableVQ,重新审视每个模块的合适学习目标,解决各模块独立训练以承担各自角色时产生的问题;其基于共享投影 codebook 构建,轻量且不引入可学习参数。StableVQ is proposed, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently, and Built on top of shared-projection codebooks, is lightweight and introduces no learnable parameters.

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes
三维场景中交互理解的几何与语义耦合
arXiv:2609.25247 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了 SEGMENT-SNAP,通过 part-handle coupling 融合几何与语义证据,在 Articulate3D Challenge 中获得第一名。This work presents SEGMENT-SNAP, which combines geometric and semantic evidence through part-handle coupling and achieved first place in the Articulate3D Challenge.

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing
PackLab:面向机器人装箱场景中 MLLM 开发、训练与评估的综合框架
arXiv:2609.23784 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

PackLab 是一个用于开发、训练与评估闭环机器人装箱 MLLM 的综合框架,在不同物体集合与容器配置下均优于传统装箱启发式方法、经典强化学习方法以及通用 MLLM,展现了 MLLM 在长时任务机器人装箱中的潜力。PackLab is a comprehensive framework for developing, training, and evaluating MLLMs for closed-loop robotic bin packing that outperforms conventional packing heuristics, traditional reinforcement learning methods, and general-purpose MLLMs across object sets and container configurations, highlighting the potential of MLLMs for long-horizon robotic packing.

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
GeoPair:面向免训练 Transformer 压缩的几何保持跨层分解
arXiv:2609.25963 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个基于原理的、无需训练的框架,依次优化跨层权重配对与共享字典分解,并识别结构兼容的投影、学习一种能更好保留各层独立校准几何的共享表征。This work introduces a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations, and identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry.

Rufus-Air: An Open LLM Post-Training Recipe
Rufus-Air:一种开放的大语言模型后训练方案
arXiv:2609.29421 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

主要发现是:多样且高质量的 SFT 奠定坚实的能力下限,难度过滤将 RL 提示维持在有效的学习区间内,而奖励可靠性为阶段排序提供了实用原则。The main findings are that diverse, high-quality SFT establishes a strong capability floor and difficulty filtering keeps RL prompts within a productive learning range, and reward reliability provides a practical principle for ordering stages.

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
你的 Transformer 可同时容纳两个思考:LLM 中线性叠加的证据
arXiv:2609.29845 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文给出证据,说明 superposition 是 Transformer 架构的内在属性,而非训练过程中涌现的结果,并证明通过轻量级 fine-tuning 可以在很大程度上恢复线性性。This work provides evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training, and demonstrates that linearity can be substantially restored through lightweight fine-tuning.

CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation
[标题中文] CARD:面向个性化文本生成的基于聚类级自适应与奖励引导解码
arXiv:2601.06352 工程化 应用落地 OA · 绿色 被引 2 · S2

本文提出 CARD,一种通过渐进式细化实现有效个性化的层级框架:先按共享风格模式对用户聚类,再为各组学习专用的 LoRA adapter,从而在低资源场景下也能实现稳健的泛化与强劲的性能。This work presents CARD, a hierarchical framework that achieves effective personalization through progressive refinement that first clusters users according to shared stylistic patterns and learns group-specific LoRA adapters, enabling robust generalization and strong low-resource performance.

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
[标题中文] MOPD-Router:重新思考多教师在策略蒸馏中的教师路由
arXiv:2609.30837 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ExpertAlign,一种无需域标签或独立 routing 模型的框架,可在每个 token 上对全体 teacher 池进行监督路由;研究表明 token 级路由能够利用跨域互补监督,并减少对 prompt 级域指派的单一依赖。This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment.

Residual Transferability in Neural Image Watermarking
神经图像水印中的残差可迁移性
arXiv:2609.32241 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文识别出两种强化水印证据对载体图像依赖性、抑制残余可迁移性的机制,并提出 CoverLock,一种即插即用策略,可在不重新设计架构的前提下强化现有水印系统的图像依赖性。This work identifies two mechanisms that strengthen the dependence of watermark evidence on the cover image, thereby suppressing the residual transferability, and introduces CoverLock, a plug-and-play strategy for existing watermarking systems that strengthens such image dependence without architectural redesign.

Nereus: Adaptive Parallelism for LLM Post-Training
Nereus: 面向 LLM 后训练的自适应并行
arXiv:2609.34645 工程化 方法 OA · 绿色 被引 1 · S2

Nereus 是一种成本感知的运行时,将 RL 后训练任务适配为高效执行计划,并基于内存可行的全局计划执行状态转移,使用与运行任务校准的成本模型来接纳转移。Nereus is a cost-aware runtime that adapts RL post-training jobs into efficient execution plans and executes a transition using a memory-feasible global plan and admits the transition using a cost model calibrated against the running job.

SANTA++: Sampling Attention through Representative Keys
SANTA++: 通过代表性 Key 进行采样的注意力机制
arXiv:2609.35629 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

注意力往往集中在上下文中一小部分 token 上,但每个 query 关注的关键子集各不相同。为利用这种动态结构,我们提出 SANTA++,一种免训练的随机注意力方法,通过代表性 key 进行内存高效的选择,无需扫描整个 KV cache。缓存的 key 被组织成若干 team,query 对每个 team 中的代表性 key 打分以决定采样哪些 team。我们在采样得到的 team 内计算精确的注意力分数,并通过其采样概率的倒数对各 team 的贡献进行重新加权。Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusio

Selecting Diverse SFT Traces Improves Post-RL Generalization
选择多样化的 SFT 轨迹可提升 RL 后的泛化能力
arXiv:2609.33780 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,推理路径多样性可作为筛选 SFT 数据的实用准则,能更好地为 RL 准备模型,并据此提出一种轻量级、基于规则的指纹方法用于筛选。These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL, and propose a lightweight, rule-based fingerprint to select for it.

Pretraining Transformers with Quantized Softmax in Attention
[标题中文] 使用量化 Softmax 的注意力机制预训练 Transformer
arXiv:2609.33591 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文推导了对应的反向传播规则(包括校准项的导数),并在模型、数据、优化器均一致的预训练实验中对不同选择进行了比较。This work derives the corresponding backward rules, including calibration derivatives, and compares these choices in pretraining experiments matched on model, data, and optimizer, and compares these choices in pretraining experiments matched on model, data, and optimizer.

Scaling Properties of Same-Family On-Policy Distillation
同家族同策略蒸馏的缩放特性
arXiv:2609.32722 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现,早期 OPD 训练动态均呈现一种规律的 *useful-transfer* 区间,其中留出准确率(即 *gold score*, $G$)随 $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$(学生初始化在 token 级反向 KL 散度的平方根)近似线性上升。It is found that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization.

ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
ATLAS:面向可靠世界模型规划的对齐潜在结构传输
arXiv:2609.36333 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ATLAS(Aligned Transport of Latent Structure),一种在显式保持关系几何结构的同时校准全局潜空间分布的训练目标,通过一维 Wasserstein-2 传输进行 Wasserstein 嵌入匹配来校准其边缘分布。This work introduces Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution and uses Wasserstein embedding matching to calibrate its marginal through one-dimensional Wasserstein-2 transport.

Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
通过位置选择性自蒸馏从语言反馈训练 LLM 评判器
arXiv:2609.38792 工程化 观点 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,屏蔽高熵偏移位置相比朴素的 SD 提升了分布外泛化能力,由此得到的自蒸馏 judges 在所评估的主观子类别上比基于结果监督 RL 训练的 judges 高出 2-9 个百分点,同时在客观子类别上保持竞争力。Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD, and the resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.

Activation Functions: Comparison of trends in Practice and Research for Deep Learning
激活函数:深度学习实践与研究趋势对比
arXiv:1811.03378 工程化 综述 OA · 绿色 被引 1486 · S2

本文首次系统梳理了深度学习研究迄今为止的激活函数应用趋势,将实践中的使用情况与文献中的研究成果进行对比。This paper will be the first, to compile the trends in AF applications in practice against the research results from literature, found in deep learning research to date.

Lower bounds for multivariate independence polynomials and their generalisations
多元独立多项式及其推广的下界
arXiv:2602.02450 工程化 方法 OA · 绿色 被引 12 · S2

在统计物理中,多元硬核模型描述一个粒子系统,每个粒子拥有各自的逸度。用图论语言表述,该模型的配分函数对应多元独立多项式,即独立多项式的多重仿射推广,定义为 $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$,其中 $\mathcal{I}(G)$ 表示 $[n]:=\{1,2,\dots,n\}$ 上图 $G$ 的所有独立集。我们证明对于 $[n]$ 上的每个简单图 $G$ 以及 $λ_1,\dots,λ_n\geq 0$,\[ Z_G(λ_1,\dots,In statistical physics, the multivariate hard-core model describes a system of particles, each of which receives its own fugacity. In graph-theoretic language, the partition function of the model translates to the multivariate independence polynomial, i.e., the multiaffine generalisation of the independence polynomial, defined by $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$, where $\mathcal{I}(G)$ denotes the set of all independent sets in a graph $G$ on $[n]:=\{1,2,\dots,n\}$. We prove that for every simple graph $G$ on $[n]$ and $λ_1,\dots,λ_n\geq 0$, \[ Z_G(λ_1,\dots,

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Transformer 过早停止思考,一个微型 LoRA 即可修复
arXiv:2609.36585 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

预训练 Transformer 仅利用其深度的一小部分来跟踪上下文中的引用。十三个基础模型仅能可靠地跟随 1.4–3.6 行,额外的预训练循环收益甚微。在一个早期层上训练的 rank-8 LoRA 在所有模型权重冻结的情况下扩展了这一计算能力。Qwen3-8B 在 24 行链上的精确准确率从 15.5% 提升至 99%;更长训练的 LoRA 可达 50 行。Ouro-1.4B 经过四轮循环达到 60 行,八轮后至少达到 160 行。该 LoRA 启动了一场接力:程序行通过中间层的一段短距离传递其链身份。冻结的 head 逐层读取渐进式进展信号……Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressi