{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/project-codenet-a-large-scale-ai-for-code","title":"CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks","arxiv_id":"2105.12655","date":"2021-05-25","proceeding":null,"authors":["Ruchir Puri","David S. Kung","Geert Janssen","Wei zhang","Giacomo Domeniconi","Vladimir Zolotov","Julian Dolby","Jie Chen","Mihir Choudhury","Lindsey Decker","Veronika Thost","Luca Buratti","Saurabh Pujar","Shyam Ramji","Ulrich Finkler","Susan Malaika","Frederick Reiss"],"abstract":"Over the last several decades, software has been woven into the fabric of every aspect of our society. As software development surges and code infrastructure of enterprise applications ages, it is now more critical than ever to increase software development productivity and modernize legacy applications. Advances in deep learning and machine learning algorithms have enabled numerous breakthroughs, motivating researchers to leverage AI techniques to improve software development efficiency. Thus, the fast-emerging research area of AI for Code has garnered new interest and gathered momentum. In this paper, we present a large-scale dataset CodeNet, consisting of over 14 million code samples and about 500 million lines of code in 55 different programming languages, which is aimed at teaching AI to code. In addition to its large scale, CodeNet has a rich set of high-quality annotations to benchmark and help accelerate research in AI techniques for a variety of critical coding tasks, including code similarity and classification, code translation between a large variety of programming languages, and code performance (runtime and memory) improvement techniques. Additionally, CodeNet provides sample input and output test sets for 98.5% of the code samples, which can be used as an oracle for determining code correctness and potentially guide reinforcement learning for code quality improvements. As a usability feature, we provide several pre-processing tools in CodeNet to transform source code into representations that can be readily used as inputs into machine learning models. Results of code classification and code similarity experiments using the CodeNet dataset are provided as a reference. We hope that the scale, diversity and rich, high-quality annotations of CodeNet will offer unprecedented research opportunities at the intersection of AI and Software Engineering.","url_abs":"https://arxiv.org/abs/2105.12655v2","url_pdf":"https://arxiv.org/pdf/2105.12655v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"project-codenet-a-large-scale-ai-for-code","repo_url":"https://github.com/IBM/Project_CodeNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"code-classification","task_name":"Code Classification"},{"task_slug":"code-translation","task_name":"Code Translation"},{"task_slug":"diversity","task_name":"Diversity"}],"methods":[],"datasets_introduced":[{"slug":"project-codenet","name":"Project CodeNet","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2105.12655","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2105.12655"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/IBM/Project_CodeNet","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":9},"by_repo_kind":{"official":{"samples":9,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"0bcee4eb7018f213","entry":"augment_edge","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"model-experiments/gnn-based-experiments/src/utils.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/model-experiments/gnn-based-experiments/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"0bcee4eb7018f213"}},{"code_sha256_prefix":"184dc34775d243e0","entry":"get_pos_encoding_matrix","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"model-experiments/masked-language-model/pos_enc.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/model-experiments/masked-language-model/pos_enc.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"184dc34775d243e0"}},{"code_sha256_prefix":"a79df12da738da95","entry":"loadLabels","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"Contest/ExampleSimAnalysis/TestSetEval.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/Contest/ExampleSimAnalysis/TestSetEval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a79df12da738da95"}},{"code_sha256_prefix":"2a6c452ca6abc754","entry":"load_checkpoint","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"model-experiments/gnn-based-experiments/src/utils_file.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/model-experiments/gnn-based-experiments/src/utils_file.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2a6c452ca6abc754"}},{"code_sha256_prefix":"94de0ec6531dda49","entry":"load_checkpoint_results","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"model-experiments/gnn-based-experiments/src/utils_file.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/model-experiments/gnn-based-experiments/src/utils_file.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"94de0ec6531dda49"}},{"code_sha256_prefix":"60c85e9c72ac92a0","entry":"makeDataset","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"Contest/ExampleSimAnalysis/TestSetEval.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/Contest/ExampleSimAnalysis/TestSetEval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"60c85e9c72ac92a0"}},{"code_sha256_prefix":"3639e3692a982ac9","entry":"summary_report","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"model-experiments/gnn-based-experiments/src/utils_file.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/model-experiments/gnn-based-experiments/src/utils_file.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3639e3692a982ac9"}},{"code_sha256_prefix":"2eea921b0bf4fe00","entry":"tokenize","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"model-experiments/masked-language-model/infer.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/model-experiments/masked-language-model/infer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2eea921b0bf4fe00"}},{"code_sha256_prefix":"4b2e20f1da5bb1d8","entry":"tokenizeFile","repo":"IBM/Project_CodeNet","repo_kind":"official","path":"Contest/ExampleSimAnalysis/TestSetEval.py","file_url":"https://github.com/IBM/Project_CodeNet/blob/HEAD/Contest/ExampleSimAnalysis/TestSetEval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"4b2e20f1da5bb1d8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}