MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

  • 类型:arxiv
  • 标识:2609.10016
  • 链接:https://arxiv.org/abs/2609.10016
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a
  • 待LLM分类:否
  • 标题中文:MetroLLM-Bench:将语言模型作为交通售票机运行时进行评估
  • TLDR中文:我们推出 MetroLLM-Bench,这是一个包含 955 个用例的基准,用于测试语言模型作为交通售票机策略层的能力。它涵盖六个真实地铁系统,规模从 37 个站点到 414 个站点不等,包含十一个类别,涉及路径规划、票价计算、运行中断、无障碍服务以及对抗性输入。在每个用例中,模型必须调用结构化工具并提交一个机器可渲染的终态,其中包含结果、适用的每张票票价报价以及售票机动作。四个确定性评分组件构成 Tier 1;八个语义质量组件构成 Tier 2,其中六个使用
  • 来源文件
  • /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json