{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-language-models-for-code-syntax","title":"Benchmarking Language Models for Code Syntax Understanding","arxiv_id":"2210.14473","date":"2022-10-26","proceeding":null,"authors":["Da Shen","Xinyun Chen","Chenguang Wang","Koushik Sen","Dawn Song"],"abstract":"Pre-trained language models have demonstrated impressive performance in both natural language processing and program understanding, which represent the input as a token sequence without explicitly modeling its structure. Some prior works show that pre-trained language models can capture the syntactic rules of natural languages without finetuning on syntax understanding tasks. However, there is limited understanding of how well pre-trained models understand the code structure so far. In this work, we perform the first thorough benchmarking of the state-of-the-art pre-trained models for identifying the syntactic structures of programs. Specifically, we introduce CodeSyntax, a large-scale dataset of programs annotated with the syntactic relationships in their corresponding abstract syntax trees. Our key observation is that existing language models pretrained on code still lack the understanding of code syntax. In fact, these pre-trained programming language models fail to match the performance of simple baselines based on positional offsets and keywords. We also present a natural language benchmark to highlight the differences between natural languages and programming languages in terms of syntactic structure understanding. Our findings point out key limitations of existing pre-training methods for programming languages, and suggest the importance of modeling code syntactic structures.","url_abs":"https://arxiv.org/abs/2210.14473v1","url_pdf":"https://arxiv.org/pdf/2210.14473v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-language-models-for-code-syntax","repo_url":"https://github.com/dashends/codesyntax","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[{"method_slug":"fail","method_name":"fail"}],"datasets_introduced":[{"slug":"codesyntax","name":"CodeSyntax","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.14473","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2210.14473"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dashends/codesyntax","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"unverified":5},"by_repo_kind":{"official":{"samples":6,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"70f4d7f0b484bff1","entry":"load_pickle","repo":"dashends/codesyntax","repo_kind":"official","path":"evaluating_models/NL/preprocess_attn_NL_word_level_sorted.py","file_url":"https://github.com/dashends/codesyntax/blob/HEAD/evaluating_models/NL/preprocess_attn_NL_word_level_sorted.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"70f4d7f0b484bff1"}},{"code_sha256_prefix":"5430d736b8cf5809","entry":"align_codebert_tokens","repo":"dashends/codesyntax","repo_kind":"official","path":"evaluating_models/NL/run_exp_roberta.py","file_url":"https://github.com/dashends/codesyntax/blob/HEAD/evaluating_models/NL/run_exp_roberta.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5430d736b8cf5809"}},{"code_sha256_prefix":"5d9b89c5a3f31f33","entry":"convert_line_offset_to_char_offset_java","repo":"dashends/codesyntax","repo_kind":"official","path":"generating_CodeSyntax/generate_labels_java.py","file_url":"https://github.com/dashends/codesyntax/blob/HEAD/generating_CodeSyntax/generate_labels_java.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5d9b89c5a3f31f33"}},{"code_sha256_prefix":"8d00fe0914d77239","entry":"get_label","repo":"dashends/codesyntax","repo_kind":"official","path":"generating_CodeSyntax/generate_labels_python.py","file_url":"https://github.com/dashends/codesyntax/blob/HEAD/generating_CodeSyntax/generate_labels_python.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8d00fe0914d77239"}},{"code_sha256_prefix":"3a8052c9a94959b8","entry":"get_word_word_attention","repo":"dashends/codesyntax","repo_kind":"official","path":"evaluating_models/NL/run_exp_roberta.py","file_url":"https://github.com/dashends/codesyntax/blob/HEAD/evaluating_models/NL/run_exp_roberta.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3a8052c9a94959b8"}},{"code_sha256_prefix":"5ea0facdf9ebd883","entry":"make_attn_word_level","repo":"dashends/codesyntax","repo_kind":"official","path":"evaluating_models/NL/run_exp_roberta.py","file_url":"https://github.com/dashends/codesyntax/blob/HEAD/evaluating_models/NL/run_exp_roberta.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5ea0facdf9ebd883"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}