Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

  • 类型:arxiv
  • 标识:2305.01210
  • 链接:https://arxiv.org/abs/2305.01210
  • 主题:engineering
  • 主分类:evaluation
  • 形态:benchmark
  • 被引:2078
  • 被引来源:Semantic Scholar
  • S2被引:2078
  • OpenAlex被引:177
  • 影响力被引:200
  • TLDR:EvalPlus -- a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code and augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM and mutation-based strategies.
  • OpenAlex ID:W4367860052
  • OpenAlex DOI:10.48550/arxiv.2305.01210
  • DOI:10.48550/arxiv.2305.01210
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/2305.01210
  • OpenAlex更新:2026-08-18
  • 待LLM分类:否
  • 标题中文:ChatGPT 生成的代码真的正确吗?面向代码生成的大型语言模型严格评估
  • TLDR中文:EvalPlus——一个用于严格基准测试 LLM 生成代码功能正确性的代码合成评估框架,通过 LLM 与基于 mutation 的策略驱动的自动测试输入生成器,为给定评估数据集补充大量新生成的测试用例。
  • 来源文件
  • [OpenAlex discover]
  • [S2 enrich]