Calibration as a First-Class Criterion in LLM Evaluation
- 类型:arxiv
- 标识:2609.26489
- 链接:https://arxiv.org/abs/2609.26489
- 主分类:evaluation
- 形态:benchmark
- TLDR:Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the res
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-24-agent-rag-longcontext-candidates.json