FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

  • 类型:arxiv
  • 标识:2608.20574
  • 链接:https://arxiv.org/abs/2608.20574
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and fam
  • 待LLM分类:否
  • 标题中文:FlavourBench:基于可执行烹饪真值的 frontier 语言模型排名
  • TLDR中文:开放式语言模型基准通常继承一种评判方式:人类偏好面板、另一个模型,或脆弱的精确匹配答案。我们提出 FlavourBench,一个自动化基准,其中版本化的烹饪系统提供密集、可执行的真值。每个任务给出八种食材,要求选择三食材组合;在模型执行前,Epicure 对全部 56 种可能组合打分。我们在相同的核心任务集上评估了 27 个 frontier 端点,覆盖替换、配对与受限组合共 534 道任务。每个被排名的模型在每个面板上恰好有 89 个有效回答,且 fam
  • 来源文件
  • /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json
  • /inbox/tom/_candidates/2026-08-25-agent-memory-tool-use-candidates.json