StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- 类型:arxiv
- 标识:2608.15089
- 链接:https://arxiv.org/abs/2608.15089
- 主分类:agent
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:StateM is introduced, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together to bet on harness scaling to improve the execution system around an agent without changing its model weights.
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- TLDR中文:提出 StateM,一种以持久化 state、phase-local 上下文、受检 transition、可恢复 runbook 以及版本化流程实践为核心组织执行的 agent-native runtime,使 Agent 与用户能够共同检视其执行过程,从而在不改模型权重的前提下通过 harness scaling 改善 Agent 周边的执行系统。
- 来源文件:
- /inbox/tom/_candidates/2026-08-19-agent-rag-longcontext-candidates.json
- [S2 enrich]