Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
- 类型:arxiv
- 标识:2608.25417
- 链接:https://arxiv.org/abs/2608.25417
- 主分类:agent
- 形态:benchmark
- TLDR:Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of
- 副分类:evaluation
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json