Datasets › HumanEval
HumanEval
This is an evaluation harness for the HumanEval problem solving dataset described in the paper "Evaluating Large Language Models Trained on Code". It used to measure functional correctness for synthesizing programs from docstrings. It consists of 164 original programming problems, assessing language comprehension, algorithms, and simple mathematics, with some comparable to simple software interview questions.
Source: Evaluating Large Language Models Trained on Code Image Source: Evaluating Large Language Models Trained on Code
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Code Generation | HumanEval | DeepSeek-R1 (MGDebugger) Pass@1 100 | From Code to Correctness: Closing the Last Mile of Code... | YerbaPage/MGDebugger | 8 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,201. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Execution Guided Line-by-Line Code Generation | 1 | 1 | 12 Jun 2025 | not harvested |
| QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks | 0 | 1 | 20 Jan 2025 | not harvested |
| Planning-Driven Programming: A Large Language Model Programming Workflow | 1 | 1 | 21 Nov 2024 | ran 0 of 9 samples (9 unverified) |
| From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging | 1 | 1 | 2 Oct 2024 | ran 8 of 9 samples (1 unverified) |
| MapCoder: Multi-Agent Code Generation for Competitive Problem Solving | 2 | 1 | 18 May 2024 | ran 0 of 10 samples (10 unverified) |
| Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step | 1 | 1 | 25 Feb 2024 | ran 8 of 12 samples (4 unverified) |
| L2MAC: Large Language Model Automatic Computer for Extensive Code Generation | 2 | 1 | 2 Oct 2023 | ran 10 of 14 samples (4 unverified) |
Dataset loaders archive 2025-07-28
3 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Unknown
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- STS Benchmark
- HumanEval
- HumanEval!
- humaneval (0-shots)
- MTEB Benchmark
- AllNLI Triplet
6 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections