Graph Machine: Towards Better Pretraining via Edges

  • 类型:arxiv
  • 标识:2609.02881
  • 链接:https://arxiv.org/abs/2609.02881
  • 主分类:engineering
  • 形态:method
  • TLDR:We introduce the Graph Machine (GM), an architecture that maintains an O(n)-sized state and accesses it through sparse, dynamic routing. Unlike methods with fixed-size states or sparse but static routing, GM preserves O(n) complexity in its sparse layers without restricting the potentially accessible state size to O(1). Instead, GM uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing. We replace 75% of the dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrain from scratch on 15.7B tokens. With only 2 of 4,096 tokens ret
  • 待LLM分类:否
  • 标题中文:Graph Machine:通过边实现更优的预训练
  • TLDR中文:我们提出 Graph Machine(GM),一种维持 O(n) 规模状态、并通过稀疏动态路由访问该状态的架构。与采用固定大小状态或稀疏但静态路由的方法不同,GM 在稀疏层中保持 O(n) 复杂度,且不把潜在可访问的状态大小限制为 O(1)。GM 使用边——一类类似指针的对象,通过类似指针追踪的引用机制以可微分方式更新。我们将 Qwen3-0.6B 中 75% 的密集 Transformer 层替换为 GM 稀疏层,并在 15.7B token 上从头预训练。在 4,096 个 token 中仅 2 个被
  • 来源文件
  • /inbox/tom/_candidates/2026-09-10-agent-rag-longcontext-candidates.json