Parameter Exploration for RLVR via Variational Learning
- 类型:arxiv
- 标识:2608.09805
- 链接:https://arxiv.org/abs/2608.09805
- 主分类:engineering
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:Evidence that parameter-space exploration can improve reinforcement learning for LLMs is presented, and a family of methods called Perturbed Parameter Policy Optimization (3PO) is introduced which use different sampling strategies and different rollout grouping for reward estimation.
- 待LLM分类:否
- 标题中文:通过变分学习实现 RLVR 的参数空间探索
- TLDR中文:本文提供了参数空间探索能够改进 LLM 强化学习的证据,并提出称为扰动参数策略优化(3PO)的方法族,使用不同的采样策略和不同的 rollout 分组进行 reward 估计。
- 来源文件:
- /inbox/tom/_candidates/2026-08-13-agent-rag-longcontext-candidates.json
- /inbox/tom/_candidates/2026-08-14-agent-rag-longcontext-candidates.json
- [S2 enrich]