Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
- 类型:arxiv
- 标识:2609.22220
- 链接:https://arxiv.org/abs/2609.22220
- 主分类:evaluation
- 形态:benchmark
- TLDR:Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuzzing recipes, tighter tolerances---with no way to measure whether any patch suffices. We introduce mutation analysis as an adequacy metric for kernel-benchmark oracles: deterministic rules inject 10{,}303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of th
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-22-agent-rag-longcontext-candidates.json