FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
- 类型:arxiv
- 标识:2608.20574
- 链接:https://arxiv.org/abs/2608.20574
- 主分类:evaluation
- 形态:benchmark
- TLDR:Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and fam
- 待LLM分类:否
- 标题中文:FlavourBench:基于可执行烹饪真值的 frontier 语言模型排名
- TLDR中文:开放式语言模型基准通常继承一种评判方式:人类偏好面板、另一个模型,或脆弱的精确匹配答案。我们提出 FlavourBench,一个自动化基准,其中版本化的烹饪系统提供密集、可执行的真值。每个任务给出八种食材,要求选择三食材组合;在模型执行前,Epicure 对全部 56 种可能组合打分。我们在相同的核心任务集上评估了 27 个 frontier 端点,覆盖替换、配对与受限组合共 534 道任务。每个被排名的模型在每个面板上恰好有 89 个有效回答,且 fam
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json
- /inbox/tom/_candidates/2026-08-25-agent-memory-tool-use-candidates.json