Nereus: Adaptive Parallelism for LLM Post-Training

  • 类型:arxiv
  • 标识:2609.34645
  • 链接:https://arxiv.org/abs/2609.34645
  • 主分类:engineering
  • 形态:method
  • TLDR:Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating
  • 待LLM分类:否
  • 标题中文:Nereus: 面向 LLM 后训练的自适应并行
  • TLDR中文:大语言模型(LLM)的强化学习(RL)后训练在 GPU 集群上协调生成、推理与训练中的多个模型。运行过程中资源可用性、序列长度、内存压力与阶段瓶颈等因素可能发生变化,使原本合适的执行计划随时间变慢甚至不可行。然而,调整一个模型共享 GPU 的作业面临重大挑战:判断新计划是否值得迁移成本、复用作业的分布式状态、协调
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json