Nereus: Adaptive Parallelism for LLM Post-Training
- 类型:arxiv
- 标识:2609.34645
- 链接:https://arxiv.org/abs/2609.34645
- 主分类:engineering
- 形态:method
- TLDR:Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating
- 待LLM分类:否
- 标题中文:Nereus: 面向 LLM 后训练的自适应并行
- TLDR中文:大语言模型(LLM)的强化学习(RL)后训练在 GPU 集群上协调生成、推理与训练中的多个模型。运行过程中资源可用性、序列长度、内存压力与阶段瓶颈等因素可能发生变化,使原本合适的执行计划随时间变慢甚至不可行。然而,调整一个模型共享 GPU 的作业面临重大挑战:判断新计划是否值得迁移成本、复用作业的分布式状态、协调
- 来源文件:
- /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json