SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

  • 类型:arxiv
  • 标识:2610.02304
  • 链接:https://arxiv.org/abs/2610.02304
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:Existing Simulink benchmarks mainly evaluate whether generated models compile, execute, or resemble a reference model. These criteria do not establish whether a model satisfies its engineering requirements. We introduce SimuVerity, a benchmark of 101 text-to-executable Simulink model-generation tasks across ten engineering domains. For each task, executable-system profiles ground the engineering specification and four families of native simulation scenarios. A hierarchical evaluator first checks artifact delivery, native executability, and engineering qualification, then scores qualified model
  • 副分类:agent
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json