Home › Datasets › task › Code Generation

Code Generation datasets

archive 2025-07-28

70 datasets carry the task tag "Code Generation" (the task itself: Code Generation), ordered by the archive's paper count. Page 2 of 2: 22 shown of 70. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Code Generation datasets 49–70 of 70

Educational Grade School Math (EGSM) contains 2,093 question/answer pairs generated by MATHWELL, a reference-free educational grade school math word problem generator that outputs a word problem and Program of Thought (PoT) solution based…
1 paper · 0 benchmarks
This dataset contains: (1) Slforge Generated Simulink Models : Synthetic Simulink Models (2) Source of Real World Simulink Models The .txt file is a combined text file that contains all the real world Simulink models based on SLGPT's…
1 paper · 0 benchmarks
FloCo (Flow chart Image to Code)
the FloCo dataset that contains 11,884 flowchart images and their corresponding Python codes.
1 paper · 1 benchmark
GenoTEX (An LLM Agent Benchmark for Automated Gene Expression Data Analysis)
GenoTEX (Genomics Data Automatic Exploration Benchmark) is a benchmark dataset for the automated analysis of gene expression data to identify disease-associated genes while considering the influence of other biological factors.
1 paper · 0 benchmarks
Source code to obfuscated code dataset in C, C++, Go, Java, Python, Rust and TypeScript.
1 paper · 0 benchmarks
A human-refined dataset of OpenAPI definitions based on the APIs.guru OpenAPI directory.
1 paper · 1 benchmark
PECC (PECC: Problem Extraction and Coding Challenges)
Recent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
RES-Q (RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository Scale)
RES-Q is a natural language instruction-based benchmark for evaluating Repository Editing Systems, which consists of 100 handcrafted repository editing tasks derived from real GitHub commits.
1 paper · 1 benchmark
SOEVAL is created by us by mining questions from StackOverflow.
1 paper · 0 benchmarks
TACO-BAAI (Topics in Algorithmic Code generation dataset)
TACO (Topics in Algorithmic Code generation dataset) is a dataset focused on algorithmic code generation, designed to provide a more challenging training dataset and evaluation benchmark for the code generation model field.
1 paper · 1 benchmark
The dataset contains more than 100k code patch pairs extracted from open source projects on GitHub.
1 paper · 1 benchmark
Turbulence is a new benchmark for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation.
1 paper · 1 benchmark
Verified Smart Contracts Code Comments is a dataset of real Ethereum smart contract functions, containing "code, comment" pairs of both Solidity and Vyper source code.
1 paper · 1 benchmark
Verified Smart Contracts is a dataset of real Ethereum smart contracts, containing both Solidity and Vyper source code.
1 paper · 0 benchmarks
Verireason-RTL-Coder_7b_reasoning_tb (VeriReason Verilog Dataset with Reasoning, Testbench, and Simulation Results)
Verireason-RTL-Coder7breasoningtb For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log…
1 paper · 0 benchmarks
Verireason-RTL-Coder_7b_reasoning_tb_simple (Simple Problems of VeriReason Verilog Dataset with Reasoning, Testbench, and Simulation Results)
Verireason-RTL-Coder7breasoningtbsimple For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update…
1 paper · 0 benchmarks
Test-driven benchmark to challenge LLMs to write long JavaScript React application GitHub Script
1 paper · 1 benchmark
WebGen-Bench WebGen-Bench is created to benchmark LLM-based agent's ability to generate websites from scratch.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.