{"url":"/method/cubert","slug":"cubert","name":"CuBERT","full_name":"CuBERT","full_name_withheld":false,"description_markdown":"**CuBERT**, or **Code Understanding BERT**, is a [BERT](https://paperswithcode.com/method/bert) based model for code understanding. In order to achieve this, the authors curate a massive corpus of Python programs collected from GitHub. GitHub projects are known to contain a large amount of duplicate code. To avoid biasing the model to such duplicated code, authors perform deduplication using the method of [Allamanis (2018)](https://arxiv.org/abs/1812.06469). The resulting corpus has 7.4 million files with a total of 9.3 billion tokens (16 million unique).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Learning and Evaluating Contextual Embedding of Source Code","paper":"/paper/pre-trained-contextual-embedding-of-source-1","first_author":"Aditya Kanade","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/pre-trained-contextual-embedding-of-source-1"},"source":{"url":"https://arxiv.org/abs/2001.00059v3","title":"Learning and Evaluating Contextual Embedding of Source Code","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Code Generation Transformers","url":"/methods/category/code-generation-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoencoding Transformers","url":"/methods/category/autoencoding-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Models","url":"/methods/category/language-models","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":3,"papers_newest_first":[{"paper":null,"title":"Intraoperative perfusion assessment by continuous, low-latency hyperspectral light-field imaging: development, methodology, and clinical application","date":"2025-04-15","arxiv_id":"2504.10953","n_code_links":0,"syntology":null},{"paper":null,"title":"SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation","date":"2021-08-10","arxiv_id":"2108.04556","n_code_links":0,"syntology":null},{"paper":"/paper/pre-trained-contextual-embedding-of-source-1","title":"Learning and Evaluating Contextual Embedding of Source Code","date":"2019-12-21","arxiv_id":"2001.00059","n_code_links":2,"syntology":null}],"papers_shown":3,"tasks":[{"task":"/task/clone-detection","name":"Clone Detection","papers":1},{"task":"/task/code-search","name":"Code Search","papers":1},{"task":"/task/code-translation","name":"Code Translation","papers":1},{"task":"/task/contextual-embedding-for-source-code","name":"Contextual Embedding for Source Code","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/exception-type","name":"Exception type","papers":1},{"task":"/task/function-docstring-mismatch","name":"Function-docstring mismatch","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1},{"task":"/task/program-repair","name":"Program Repair","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/swapped-operands","name":"Swapped operands","papers":1},{"task":"/task/type-prediction","name":"Type prediction","papers":1},{"task":"/task/variable-misuse","name":"Variable misuse","papers":1},{"task":"/task/wrong-binary-operator","name":"Wrong binary operator","papers":1}],"tasks_shown":15,"n_tasks":15,"usage_by_year":[{"year":"2019","papers":1},{"year":"2021","papers":1},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cubert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}