ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
- 类型:arxiv
- 标识:2609.14857
- 链接:https://arxiv.org/abs/2609.14857
- 主分类:evaluation
- 形态:benchmark
- TLDR:Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve execution mechanisms from experience. However, generalizable harness RSI remains challenging. First, evolving harnesses on evaluation benchmarks or their subsets makes it difficult to distinguish reusable improvements from benchmark-specific adaptation. Second, single-trajectory updates can conflate systematic harness deficiencies with instance-specific reasoning and solution details, producing modifications that transfer poorly to unseen tasks. Third,
- 副分类:agent
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-16-agent-rag-longcontext-candidates.json