{"url":"/dataset/codesyntax","name":"CodeSyntax","full_name":null,"description_markdown":"**CodeSyntax** is a large-scale dataset of programs annotated with the syntactic relationships in their corresponding abstract syntax trees. It contains 18,701 code samples annotated with 1,342,050 relation edges in 43 relation types for Python, and 13,711 code samples annotated with 864,411 relation edges in 39 relation types for Java. It is designed to  evaluate the performance of language models on code syntax understanding.\r\n\r\nSource: [https://paperswithcode.com/paper/benchmarking-language-models-for-code-syntax](https://arxiv.org/pdf/2210.14473v1.pdf)\r\n\r\nImage Source: [https://arxiv.org/pdf/2210.14473v1.pdf](https://arxiv.org/pdf/2210.14473v1.pdf)","description_withheld":null,"homepage":"https://github.com/dashends/CodeSyntax","introduced_date":"2022-10-26","introduced_date_note":null,"introduced_by":{"paper":"/paper/benchmarking-language-models-for-code-syntax","title":"Benchmarking Language Models for Code Syntax Understanding","first_author":"Da Shen","url":null},"license":{"name":"MIT license","url":"https://github.com/dashends/CodeSyntax/blob/main/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Generation","url":"/task/text-generation","datasets_with_task":"/datasets/task/text-generation"}],"languages":[],"variants":["CodeSyntax"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}