TLDR
大语言模型可能持有但不报告的知识。模型可能在能力评估中刻意隐藏,或给出与内部知识相反的答案,仅凭其输出无法区分是在隐藏答案还是真不知道。我们借鉴隐蔽信息测试——一种通过在多个合理干扰项中呈现真实细节并测量对识别项的更强反应来识别内隐知识的取证方法。我们的方法内部识别探测(PIR)在模型内部复现了这一过程。它向模型呈现一个问题及其候选答案Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the Concealed Information Test, a forensic method that identifies guilty knowledge by presenting a suspect with the true detail among plausible decoys and measuring a stronger response to the item they recognize. Our method, Probe of Internal Recognition (PIR), does the same inside a model. It presents a question with its candidate answers a