The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
- 类型:arxiv
- 标识:2609.25804
- 链接:https://arxiv.org/abs/2609.25804
- 主分类:agent
- 形态:benchmark
- TLDR:LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them measures the taste of an agent. To address this problem, we build Taste-Bench, a benchmark of taste questio
- 副分类:evaluation
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-23-agent-rag-longcontext-candidates.json