Datasets › HumanEval

HumanEval

Introduced by Mark Chen et al. in Evaluating Large Language Models Trained on Code7 Jul 2021 archive 2025-07-28

This is an evaluation harness for the HumanEval problem solving dataset described in the paper "Evaluating Large Language Models Trained on Code". It used to measure functional correctness for synthesizing programs from docstrings. It consists of 164 original programming problems, assessing language comprehension, algorithms, and simple mathematics, with some comparable to simple software interview questions.

Source: Evaluating Large Language Models Trained on Code Image Source: Evaluating Large Language Models Trained on Code

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Code Generation HumanEval DeepSeek-R1 (MGDebugger) Pass@1 100 From Code to Correctness: Closing the Last Mile of Code... YerbaPage/MGDebugger 8 Compare

Papers archive 2025-07-28

7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,201. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Execution Guided Line-by-Line Code Generation 1 1 12 Jun 2025 not harvested
QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks 0 1 20 Jan 2025 not harvested
Planning-Driven Programming: A Large Language Model Programming Workflow 1 1 21 Nov 2024 ran 0 of 9 samples (9 unverified)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging 1 1 2 Oct 2024 ran 8 of 9 samples (1 unverified)
MapCoder: Multi-Agent Code Generation for Competitive Problem Solving 2 1 18 May 2024 ran 0 of 10 samples (10 unverified)
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step 1 1 25 Feb 2024 ran 8 of 12 samples (4 unverified)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation 2 1 2 Oct 2023 ran 10 of 14 samples (4 unverified)

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • STS Benchmark
  • HumanEval
  • HumanEval!
  • humaneval (0-shots)
  • MTEB Benchmark
  • AllNLI Triplet

6 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections