Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- 类型:arxiv
- 标识:2305.01210
- 链接:https://arxiv.org/abs/2305.01210
- 主题:engineering
- 主分类:evaluation
- 形态:benchmark
- 被引:2078
- 被引来源:Semantic Scholar
- S2被引:2078
- OpenAlex被引:177
- 影响力被引:200
- TLDR:EvalPlus -- a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code and augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM and mutation-based strategies.
- OpenAlex ID:W4367860052
- OpenAlex DOI:10.48550/arxiv.2305.01210
- DOI:10.48550/arxiv.2305.01210
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://arxiv.org/pdf/2305.01210
- OpenAlex更新:2026-08-18
- 待LLM分类:否
- 标题中文:ChatGPT 生成的代码真的正确吗?面向代码生成的大型语言模型严格评估
- TLDR中文:EvalPlus——一个用于严格基准测试 LLM 生成代码功能正确性的代码合成评估框架,通过 LLM 与基于 mutation 的策略驱动的自动测试输入生成器,为给定评估数据集补充大量新生成的测试用例。
- 来源文件:
- [OpenAlex discover]
- [S2 enrich]