LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
- 类型:arxiv
- 标识:2607.14952
- 链接:http://arxiv.org/abs/2607.14952v1
- 主分类:engineering
- 形态:application
- 被引:2
- 被引来源:Semantic Scholar
- S2被引:2
- OpenAlex被引:0
- 影响力被引:0
- TLDR:This work presents LongStraw, an objective-aware, architecture-aware system for resident-state virtualization, response replay, and distributed-gradient execution that bounds the live training graph by the response suffix while reusing the expensive prompt computation across the complete GRPO group.
- OpenAlex ID:W7169616292
- OpenAlex DOI:10.48550/arxiv.2607.14952
- DOI:10.48550/arxiv.2607.14952
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2607.14952
- OpenAlex更新:2026-07-19
- 副分类:llm-infra
- 待LLM分类:否
- 标题中文:LongStraw:固定 GPU 预算下超越 2M token 的长上下文强化学习
- TLDR中文:本文提出 LongStraw,一个面向目标、感知架构的系统,用于 resident-state 虚拟化、response replay 和分布式梯度执行,它将实时训练图限制在 response 后缀范围内,同时在完整的 GRPO 组内复用代价高昂的 prompt 计算。
- 来源文件:
- /inbox/tom/_candidates/2026-07-17-agent-rag-longcontext-candidates.json
- [S2 enrich]
- /inbox/tom/_candidates/2026-07-18-agent-rag-longcontext-candidates.json
- /inbox/tom/_candidates/2026-07-19-agent-rag-longcontext-candidates.json
- [OpenAlex backfill]
- /inbox/tom/_candidates/2026-07-20-agent-rag-longcontext-candidates.json