Home › Datasets › task › Code Generation
Code Generation datasets
archive 2025-07-28
70 datasets carry the task tag "Code Generation" (the task itself: Code Generation), ordered by the archive's paper count. Page 1 of 2: 48 shown of 70. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Code Generation datasets 1–48 of 70
This is an evaluation harness for the HumanEval problem solving dataset described in the paper "Evaluating Large Language Models Trained on Code".
1,201 papers · 1 benchmark
MBPP (Mostly Basic Python Programming)
The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry-level programmers, covering programming fundamentals, standard library functionality, and so on.
666 papers · 1 benchmark
WikiSQL consists of a corpus of 87,726 hand-annotated SQL query and natural language question pairs.
267 papers · 4 benchmarks
CodeXGLUE is a benchmark dataset and open challenge for code intelligence.
205 papers · 10 benchmarks
APPS (Automated Programming Progress Standard)
The APPS dataset consists of problems collected from different open-access coding websites such as Codeforces, Kattis, and more.
180 papers · 1 benchmark
CodeContests is a competitive programming dataset for machine-learning.
84 papers · 1 benchmark
CoNaLa (CMU CoNaLa, the Code/Natural Language Challenge)
The CMU CoNaLa, the Code/Natural Language Challenge dataset is a joint project from the Carnegie Mellon University NeuLab and Strudel labs.
77 papers · 1 benchmark
DS-1000 is a code generation benchmark with a thousand data science questions spanning seven Python libraries that (1) reflects diverse, realistic, and practical use cases, (2) has a reliable metric, (3) defends against memorization by…
72 papers · 0 benchmarks
A new large dataset with over 100,000 examples consisting of Java classes from online code repositories, and develop a new encoder-decoder architecture that models the interaction between the method documentation and the class environment.
46 papers · 1 benchmark
BigCodeBench is an easy-to-use benchmark for code generation with practical and challenging programming tasks¹.
38 papers · 2 benchmarks
In this work, we make the first attempt to evaluate LLMs in a more challenging code generation scenario, i.e.
27 papers · 0 benchmarks
The Django dataset is a dataset for code generation comprising of 16000 training, 1000 development and 1805 test annotations.
22 papers · 1 benchmark
This dataset contains card descriptions of the card game Hearthstone and the code that implements them.
22 papers · 0 benchmarks
Automated source code generation is currently a popular machine learning-based task.
17 papers · 0 benchmarks
HumanEval-X is a benchmark for evaluating the multilingual ability of code generative models.
16 papers · 0 benchmarks
JuICe is a corpus of 1.5 million examples with a curated test set of 3.7K instances based on online programming assignments.
16 papers · 0 benchmarks
Extension test cases of HumanEval, as well as generated code.
15 papers · 1 benchmark
Extension test cases of MBPP, as well as generated code.
15 papers · 1 benchmark
LLaMEA (algorithms and experiments from the paper)
3500+ Generated evolutionary algorithms by the LLaMEA framework.
12 papers · 0 benchmarks
CriticBench is a comprehensive benchmark designed to assess the abilities of Large Language Models (LLMs) to critique and rectify their reasoning across various tasks.
10 papers · 0 benchmarks
MCoNaLa is a multilingual dataset to benchmark code generation from natural language commands extending beyond English.
10 papers · 0 benchmarks
Evaluate a natural language code generation model on real data science pedagogical notebooks!
7 papers · 0 benchmarks
We introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency.
6 papers · 0 benchmarks
Lyra is a dataset for code generation that consists on Python code with embedded SQL.
6 papers · 0 benchmarks
ShellcodeIA32 is a dataset containing 20 years of shellcodes from a variety of sources is the largest collection of shellcodes in assembly available to date.
6 papers · 1 benchmark
BioCoder is a benchmark developed to evaluate existing pre-trained models in generating bioinformatics code.
5 papers · 0 benchmarks
PyTorrent contains 218,814 Python package libraries from PyPI and Anaconda environment.
5 papers · 0 benchmarks
SAFIM (Syntax-Aware Fill-In-the-Middle)
Syntax-Aware Fill-in-the-Middle (SAFIM) is a benchmark for evaluating Large Language Models (LLMs) on the code Fill-in-the-Middle (FIM) task.
5 papers · 1 benchmark
The CoNaLa Extended With Question Text is an extension to the original CoNaLa Dataset (Papers With Code Link) proposed in the NLP4Prog workshop paper "Reading StackOverflow Encourages Cheating: Adding Question Text Improves Extractive Code…
4 papers · 1 benchmark
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
4 papers · 1 benchmark
LeetCode-Hard is a benchmark dataset for code generation, consisting of 40 challenging LeetCode "hard-level" questions across 19 programming languages.
4 papers · 0 benchmarks
MMCode is a multi-modal code generation dataset designed to evaluate the problem-solving skills of code language models in visually rich contexts (i.e.
3 papers · 0 benchmarks
Description A dataset of assembly functions that are vulnerable to Spectre-V1 attack.
3 papers · 0 benchmarks
Test-driven benchmark to challenge LLMs to write JavaScript React application GitHub Script
3 papers · 1 benchmark
DISL (Fueling Research with A Large Dataset of Solidity Smart Contracts)
DISL The full dataset report is available at: https://arxiv.org/abs/2403.16861 The DISL dataset features a collection of 514, 506 unique Solidity files that have been deployed to Ethereum mainnet.
2 papers · 0 benchmarks
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
2 papers · 0 benchmarks
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
2 papers · 0 benchmarks
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
2 papers · 0 benchmarks
To automatically generate Python and assembly programs used for security exploits, we curated a large dataset for feeding NMT techniques.
2 papers · 0 benchmarks
145k natural language and PDDL problem pairs from the Blocks World, Gripper, and Floor Tile domains.
2 papers · 0 benchmarks
The first benchmark comprising 473 prompts designed to assess the ability of LLMs to resist malicious code generation.
2 papers · 0 benchmarks
The dataset consists of source code and LLVM IR pairs generated from accepted and de-duped programming contest solutions.
2 papers · 0 benchmarks
VerilogEval Dataset The VerilogEval Dataset is a benchmark specifically designed to assess the ability of large language models (LLMs) to generate syntactically correct and functionally accurate Verilog code.
2 papers · 1 benchmark
COFFE (COFFE: A Code Efficiency Benchmark for Code Generation)
COFFE COFFE is a Python benchmark for evaluating the time efficiency of LLM-generated code.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The dataset is specifically constructed for the library-oriented code generation task, which are constructed in the paper “CodeGen4Libs: A Two-Stage Approach for Library-Oriented Code Generation”.
1 paper · 0 benchmarks
This is an assembly dataset built on top of ShellcodeIA32, a dataset for automatically generating assembly from natural language descriptions that consists of 3,200 assembly instructions, commented in the English language, which were…
1 paper · 0 benchmarks
This dataset contains samples to generate Python code for security exploits.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.