Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

  • 类型:arxiv
  • 标识:2609.22220
  • 链接:https://arxiv.org/abs/2609.22220
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuzzing recipes, tighter tolerances---with no way to measure whether any patch suffices. We introduce mutation analysis as an adequacy metric for kernel-benchmark oracles: deterministic rules inject 10{,}303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of th
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-22-agent-rag-longcontext-candidates.json