Datasets › APPS

APPS (Automated Programming Progress Standard)

Introduced by Dan Hendrycks et al. in Measuring Coding Challenge Competence With APPS20 May 2021 archive 2025-07-28

The APPS dataset consists of problems collected from different open-access coding websites such as Codeforces, Kattis, and more. The APPS benchmark attempts to mirror how humans programmers are evaluated by posing coding problems in unrestricted natural language and evaluating the correctness of solutions. The problems range in difficulty from introductory to collegiate competition level and measure coding ability as well as problem-solving.

The Automated Programming Progress Standard, abbreviated APPS, consists of 10,000 coding problems in total, with 131,836 test cases for checking solutions and 232,444 ground-truth solutions written by humans. Problems can be complicated, as the average length of a problem is 293.2 words. The data are split evenly into training and test sets, with 5,000 problems each. In the test set, every problem has multiple test cases, and the average number of test cases is 21.2. Each test case is specifically designed for the corresponding problem, enabling us to rigorously evaluate program functionality.

Source: Measuring Coding Challenge Competence With APPS

Image source: Measuring Coding Challenge Competence With APPS

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Code Generation APPS LPW (GPT-4o) Introductory Pass@1 87.2 Planning-Driven Programming: A Large Language Model... you68681/lpw 18 Compare

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 180. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging 1 1 8 Feb 2025 ran 0 of 3 samples (3 unverified)
Planning-Driven Programming: A Large Language Model Programming Workflow 1 1 21 Nov 2024 ran 0 of 9 samples (9 unverified)
MapCoder: Multi-Agent Code Generation for Competitive Problem Solving 2 1 18 May 2024 ran 0 of 10 samples (10 unverified)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence 1 1 25 Jan 2024 ran 9 of 10 samples (1 unverified)
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks 1 2 26 Dec 2023 ran 4 of 8 samples (4 unverified; 8 pointer-only for licence)
CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules 1 2 13 Oct 2023 ran 3 of 3 samples (0 unverified)
CodeT: Code Generation with Generated Tests 1 2 21 Jul 2022 ran 2 of 2 samples (0 unverified)
CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning 2 4 5 Jul 2022 ran 1 of 1 samples (0 unverified)
Competition-Level Code Generation with AlphaCode 2 2 8 Feb 2022 not harvested
Evaluating Large Language Models Trained on Code 13 1 7 Jul 2021 ran 6 of 39 samples (33 unverified; 2 pointer-only for licence)
Measuring Coding Challenge Competence With APPS 3 1 20 May 2021 ran 2 of 11 samples (9 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • APPS

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections