{"url":"/dataset/naturalcodebench","name":"NaturalCodeBench","full_name":null,"description_markdown":"NaturalCodeBench (NCB) is a comprehensive code benchmark designed to mirror the complexity and variety of scenarios in real coding tasks¹². It comprises 402 high-quality problems in Python and Java, meticulously selected from an online coding service, covering 6 different domains¹². \r\n\r\nThe seed problems of NCB are cleaned from the queries in coding online services, spanning across 6 domains: Artificial Intelligence, Data Science, Algorithm and Data Structure, Front-End, Software Engineering, and System Administration¹. \r\n\r\nHere is a summary of the number of problems in each domain¹:\r\n- Software Engineering: 132 problems\r\n- Data Science: 100 problems\r\n- Algorithm and Data Structure: 95 problems\r\n- System Administration: 33 problems\r\n- Artificial Intelligence: 28 problems\r\n- Front-End: 14 problems\r\n\r\nThe development set of NCB, which contains 140 problems (70 in Python and 70 in Java), is released for research purposes¹. The data format includes a unique identifier for each question, the problem description, testcases, setup code, a reference solution, and the domain of the problem¹.\r\n\r\n(1) GitHub - THUDM/NaturalCodeBench. https://github.com/THUDM/NaturalCodeBench.\r\n(2) NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval .... https://arxiv.org/abs/2405.04520.\r\n(3) 清华、智谱AI 团队推出代码评测基准 NaturalCodeBench .... https://blog.csdn.net/www3300300/article/details/138752139.\r\n(4) NaturalCodeBench: Examining Coding Performance .... https://www.x-mol.com/paper/1788507497390862336.\r\n(5) undefined. https://doi.org/10.48550/arXiv.2405.04520.","description_withheld":null,"homepage":"https://github.com/THUDM/NaturalCodeBench","introduced_date":"2024-05-07","introduced_date_note":null,"introduced_by":{"paper":"/paper/naturalcodebench-examining-coding-performance","title":"NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts","first_author":"Shudan Zhang","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["NaturalCodeBench"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}