A Survey on Evaluation of Large Language Models
- 类型:arxiv
- 标识:2307.03109
- 链接:https://arxiv.org/abs/2307.03109
- 主题:engineering
- 主分类:evaluation
- 形态:survey
- 被引:3721
- 被引来源:Semantic Scholar
- S2被引:3721
- OpenAlex被引:202
- 影响力被引:137
- TLDR:This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate, and offers invaluable insights to researchers in the realm of LLMs evaluation.
- OpenAlex ID:W4383605161
- OpenAlex DOI:10.48550/arxiv.2307.03109
- DOI:10.48550/arxiv.2307.03109
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://arxiv.org/pdf/2307.03109
- OpenAlex更新:2026-08-18
- 待LLM分类:否
- 标题中文:大语言模型评估综述
- TLDR中文:本文对 LLM 的评估方法进行了全面综述,围绕三个关键维度展开:评估什么、在何处评估、如何评估,并为 LLM 评估领域的研究者提供了宝贵洞见。
- 来源文件:
- [OpenAlex discover]
- [S2 enrich]