Datasets › MBPP

MBPP (Mostly Basic Python Programming)

Introduced by Jacob Austin et al. in Program Synthesis with Large Language Models16 Aug 2021 archive 2025-07-28

The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry-level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Code Generation MBPP EG-CFG (DeepSeek-V3-0324) Accuracy 96.6 Execution Guided Line-by-Line Code Generation boazlavon/eg_cfg 99 Compare

Papers archive 2025-07-28

30 shown of 34 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 666. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Execution Guided Line-by-Line Code Generation 1 2 12 Jun 2025 not harvested
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging 1 1 8 Feb 2025 ran 0 of 3 samples (3 unverified)
QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks 0 1 20 Jan 2025 not harvested
Planning-Driven Programming: A Large Language Model Programming Workflow 1 1 21 Nov 2024 ran 0 of 9 samples (9 unverified)
AFlow: Automating Agentic Workflow Generation 4 1 14 Oct 2024 not harvested
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging 1 2 2 Oct 2024 ran 8 of 9 samples (1 unverified)
MapCoder: Multi-Agent Code Generation for Competitive Problem Solving 2 3 18 May 2024 ran 0 of 10 samples (10 unverified)
SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents 0 1 23 Mar 2024 not harvested
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 1 2 12 Mar 2024 not harvested
The Claude 3 Model Family: Opus, Sonnet, Haiku 0 3 4 Mar 2024 not harvested
StarCoder 2 and The Stack v2: The Next Generation 4 1 29 Feb 2024 not harvested
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence 1 8 25 Jan 2024 ran 9 of 10 samples (1 unverified)
Mixtral of Experts 6 1 8 Jan 2024 ran 5 of 5 samples (0 unverified)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks 2 2 5 Jan 2024 not harvested
AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation 1 2 20 Dec 2023 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair 1 3 16 Nov 2023 not harvested
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models 2 2 6 Oct 2023 ran 7 of 9 samples (2 unverified)
Textbooks Are All You Need II: phi-1.5 technical report 1 1 11 Sep 2023 not harvested
Code Llama: Open Foundation Models for Code 2 14 24 Aug 2023 ran 3 of 5 samples (2 unverified; 4 pointer-only for licence)
How Does Naming Affect LLMs on Code Analysis Tasks? 0 5 24 Jul 2023 not harvested
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 4 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
WizardCoder: Empowering Code Large Language Models with Evol-Instruct 4 1 14 Jun 2023 ran 3 of 7 samples (4 unverified; 1 pointer-only for licence)
PaLM 2 Technical Report 1 1 17 May 2023 not harvested
StarCoder: may the source be with you! 4 3 9 May 2023 ran 2 of 2 samples (0 unverified)
Teaching Large Language Models to Self-Debug 2 7 11 Apr 2023 ran 0 of 1 samples (1 unverified)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X 2 1 30 Mar 2023 ran 2 of 8 samples (6 unverified)
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
LEVER: Learning to Verify Language-to-Code Generation with Execution 1 1 16 Feb 2023 ran 6 of 22 samples (16 unverified)
Coder Reviewer Reranking for Code Generation 1 10 29 Nov 2022 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)

The full list of 34 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • MBPP

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections