ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
- 类型:arxiv
- 标识:2609.16816
- 链接:https://arxiv.org/abs/2609.16816
- 主分类:evaluation
- 形态:benchmark
- TLDR:Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading. Such rubrics are reliable only if they reward honest answers over adversarial answers optimized to exploit them. Yet their robustness to such optimization remains poorly understood. We isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is to acknowledge the impossibility. We introduce ImpossibleRubrics, a benchmark of 169 impossible tasks s
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-16-agent-rag-longcontext-candidates.json