Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

  • 类型:arxiv
  • 标识:2609.36585
  • 链接:https://arxiv.org/abs/2609.36585
  • 主分类:engineering
  • 形态:method
  • TLDR:Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressi
  • 待LLM分类:否
  • 标题中文:Transformer 过早停止思考,一个微型 LoRA 即可修复
  • TLDR中文:预训练 Transformer 仅利用其深度的一小部分来跟踪上下文中的引用。十三个基础模型仅能可靠地跟随 1.4–3.6 行,额外的预训练循环收益甚微。在一个早期层上训练的 rank-8 LoRA 在所有模型权重冻结的情况下扩展了这一计算能力。Qwen3-8B 在 24 行链上的精确准确率从 15.5% 提升至 99%;更长训练的 LoRA 可达 50 行。Ouro-1.4B 经过四轮循环达到 60 行,八轮后至少达到 160 行。该 LoRA 启动了一场接力:程序行通过中间层的一段短距离传递其链身份。冻结的 head 逐层读取渐进式进展信号……
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json