S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
- 类型:arxiv
- 标识:2608.31100
- 链接:https://arxiv.org/abs/2608.31100
- 主分类:evaluation
- 形态:benchmark
- TLDR:Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use that experience to improve future decisions. We introduce S\textsuperscript{3Gym}, an interactive benchmark for evaluating LLM self-improvement through three coupled capabilities: Self-Testing, Self-Judging, and Self-Improvement. S^3Gym separates permissive exploration from strict h
- 副分类:agent
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json