Methods › Natural Language Processing › Transformers › CodeBERT
CodeBERT
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
CodeBERT is a bimodal pre-trained model for programming language (PL) and natural language (NL). CodeBERT learns general-purpose representations that support downstream NL-PL applications such as natural language code search, code documentation generation, etc. CodeBERT is developed with a Transformer-based neural architecture, and is trained with a hybrid objective function that incorporates the pre-training task of replaced token detection, which is to detect plausible alternatives sampled from generators. This enables the utilization of both bimodal data of NL-PL pairs and unimodal data, where the former provides input tokens for model training while the latter helps to learn better generators.
Papers archive 2025-07-28
30 shown of 66, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair 5 Jun 2025 · 0 repositories · arXiv:2506.04987
-
LoRACode: LoRA Adapters for Code Embeddings 7 Mar 2025 · 0 repositories · arXiv:2503.05315
-
Poster: Long PHP webshell files detection based on sliding window attention 26 Feb 2025 · 1 repository · arXiv:2502.19257
-
Less is More: On the Importance of Data Quality for Unit Test Generation 20 Feb 2025 · 0 repositories · arXiv:2502.14212
-
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights 10 Feb 2025 · 0 repositories · arXiv:2502.07049
-
Repository-level Code Search with Neural Retrieval Methods 10 Feb 2025 · 1 repository · arXiv:2502.07067
-
How to Select Pre-Trained Code Models for Reuse? A Learning Perspective 7 Jan 2025 · 1 repository · arXiv:2501.03783
-
EnStack: An Ensemble Stacking Framework of Large Language Models for Enhanced Vulnerability Detection in Source Code 25 Nov 2024 · 0 repositories · arXiv:2411.16561
-
CodeSAM: Source Code Representation Learning by Infusing Self-Attention with Multi-Code-View Graphs 21 Nov 2024 · 1 repository · arXiv:2411.14611
-
Automated Vulnerability Detection Using Deep Learning Technique 29 Oct 2024 · 0 repositories · arXiv:2410.21968
-
LecPrompt: A Prompt-based Approach for Logical Error Correction with CodeBERT 10 Oct 2024 · 0 repositories · arXiv:2410.08241
-
CLNX: Bridging Code and Natural Language for C/C++ Vulnerability-Contributing Commits Identification 11 Sep 2024 · 0 repositories · arXiv:2409.07407
-
Unlearning Trojans in Large Language Models: A Comparison Between Natural Language and Source Code 22 Aug 2024 · 0 repositories · arXiv:2408.12416
-
CodeMirage: Hallucinations in Code Generated by Large Language Models 14 Aug 2024 · 0 repositories · arXiv:2408.08333
-
Enhancing Code Translation in Language Models with Few-Shot Learning via Retrieval-Augmented Generation 29 Jul 2024 · 0 repositories · arXiv:2407.19619
-
Advanced Detection of Source Code Clones via an Ensemble of Unsupervised Similarity Measures 3 May 2024 · 1 repository · arXiv:2405.02095
-
Software Vulnerability Prediction in Low-Resource Languages: An Empirical Study of CodeBERT and ChatGPT 26 Apr 2024 · 1 repository · arXiv:2404.17110
-
Enhancing Source Code Classification Effectiveness via Prompt Learning Incorporating Knowledge Features 10 Jan 2024 · 1 repository · arXiv:2401.05544
-
Source Code is a Graph, Not a Sequence: A Cross-Lingual Perspective on Code Clone Detection 27 Dec 2023 · 0 repositories · arXiv:2312.16488
-
Naturalness of Attention: Revisiting Attention in Code Language Models 22 Nov 2023 · 0 repositories · arXiv:2311.13508
-
Learning Defect Prediction from Unrealistic Data 2 Nov 2023 · 0 repositories · arXiv:2311.00931
-
Gem5Pred: Predictive Approaches For Gem5 Simulation Time 10 Oct 2023 · 0 repositories · arXiv:2310.06290
-
A Comparative Study of Filters and Deep Learning Models to predict Diabetic Retinopathy 26 Sep 2023 · 0 repositories · arXiv:2309.15216
-
XGV-BERT: Leveraging Contextualized Language Model and Graph Neural Network for Efficient Software Vulnerability Detection 26 Sep 2023 · 0 repositories · arXiv:2309.14677
-
Code quality assessment using transformers 17 Sep 2023 · 0 repositories · arXiv:2309.09264
-
Pop Quiz! Do Pre-trained Code Models Possess Knowledge of Correct API Names? 14 Sep 2023 · 0 repositories · arXiv:2309.07804
-
How Does Naming Affect LLMs on Code Analysis Tasks? 24 Jul 2023 · 0 repositories · arXiv:2307.12488
-
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs? 13 Jul 2023 · 0 repositories · arXiv:2307.06616
-
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair 21 Jun 2023 · 0 repositories · arXiv:2307.00012
-
Automatic Code Summarization via ChatGPT: How Far Are We? 22 May 2023 · 0 repositories · arXiv:2305.12865
Tasks archive 2025-07-28
20 shown of 80 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections