TLDR
强化学习(RL)已成为提升大语言模型(LLM)推理与 Agent 能力的关键技术。尽管 FP8 量化可加速 RL 训练,但在整个 FP8 RL 流程中保持稳定性仍具挑战。此前工作主要采用 TIS 等校正技术解决训练与推理的不一致问题,我们发现全流程 FP8 RL 仍存在严重的训练不稳定,表现为训练中期熵值异常飙升以及输出乱码。我们将这种不稳定归因于一个此前被忽视的原因:compoReinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-inference mismatches using correction techniques like TIS, we reveal that full-pipeline FP8 RL still suffers from severe training instability, manifesting as anomalous mid-training entropy surges and garbled outputs. We trace this instability to a previously overlooked cause: compo