Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

  • 类型:arxiv
  • 标识:2610.01428
  • 链接:https://arxiv.org/abs/2610.01428
  • 主分类:evaluation
  • 形态:benchmark
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:The Stability-Aware Generalization Objective (SAGO) is introduced, a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring.
  • 待LLM分类:否
  • 标题中文:泛化即稳定性,而非准确率:LLM 的多轴评估
  • TLDR中文:本文提出稳定性感知的泛化目标(SAGO),一个用于衡量模型在同一输入上面对不同扰动和基准时行为变化程度的评估框架,涵盖生成一致性、内部激活、置信度以及响应镜像等多个维度的变异性。
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json
  • [S2 enrich]