Methods › Natural Language Processing › Transformers › CodeT5
CodeT5
Introduced by Yue Wang et al. in CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
CodeT5 is a Transformer-based model for code understanding and generation based on the T5 architecture. It utilizes an identifier-aware pre-training objective that considers the crucial token type information (identifiers) from code. Specifically, the denoising Seq2Seq objective of T5 is extended with two identifier tagging and prediction tasks to enable the model to better leverage the token type information from programming languages, which are the identifiers assigned by developers. To improve the natural language-programming language alignment, a bimodal dual learning objective is used for a bidirectional conversion between natural language and programming language.
Papers archive 2025-07-28
30 shown of 32, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution 18 Jun 2025 · 0 repositories · arXiv:2506.17323
-
A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair 5 Jun 2025 · 0 repositories · arXiv:2506.04987
-
ShIOEnv: A CLI Behavior-Capturing Environment Enabling Grammar-Guided Command Synthesis for Dataset Curation 23 May 2025 · 1 repository · arXiv:2505.18374
-
LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming 21 May 2025 · 0 repositories · arXiv:2505.15039Syntology ran 11 of 19 samples · 8 unverified · 19 pointer-only (licence)
-
Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks 28 Apr 2025 · 0 repositories · arXiv:2504.19444
-
Enhancing Code LLM Training with Programmer Attention 19 Mar 2025 · 0 repositories · arXiv:2503.14936
-
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign 4 Feb 2025 · 0 repositories · arXiv:2502.02068
-
How to Select Pre-Trained Code Models for Reuse? A Learning Perspective 7 Jan 2025 · 1 repository · arXiv:2501.03783
-
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation 20 Dec 2024 · 0 repositories · arXiv:2412.16135
-
Generative Fuzzy System for Sequence Generation 21 Nov 2024 · 0 repositories · arXiv:2411.13867
-
Building A Coding Assistant via the Retrieval-Augmented Language Model 21 Oct 2024 · 1 repository · arXiv:2410.16229
-
Detection Made Easy: Potentials of Large Language Models for Solidity Vulnerabilities 15 Sep 2024 · 0 repositories · arXiv:2409.10574
-
VulCatch: Enhancing Binary Vulnerability Detection through CodeT5 Decompilation and KAN Advanced Feature Extraction 13 Aug 2024 · 0 repositories · arXiv:2408.07181
-
Empirical Studies of Parameter Efficient Methods for Large Language Models of Code and Knowledge Transfer to R 16 Mar 2024 · 1 repository · arXiv:2405.01553
-
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding 5 Jan 2024 · 1 repository · arXiv:2401.03003Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)
-
PerfRL: A Small Language Model Framework for Efficient Code Optimization 9 Dec 2023 · 0 repositories · arXiv:2312.05657
-
Converting Epics/Stories into Pseudocode using Transformers 8 Dec 2023 · 0 repositories · arXiv:2312.05047
-
Learning Defect Prediction from Unrealistic Data 2 Nov 2023 · 0 repositories · arXiv:2311.00931
-
Data Augmentation for Code Translation with Comparable Corpora and Multiple References 1 Nov 2023 · 1 repository · arXiv:2311.00317Syntology ran 6 of 12 samples · 6 unverified · 12 pointer-only (licence)
-
Program Repair with Minimal Edits Using CodeT5 26 Sep 2023 · 0 repositories · arXiv:2309.14760
-
RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program Repair 12 Sep 2023 · 0 repositories · arXiv:2309.06057
-
Domain Adaptation for Code Model-based Unit Test Case Generation 15 Aug 2023 · 0 repositories · arXiv:2308.08033
-
Exploring Continual Learning for Code Generation Models 5 Jul 2023 · 0 repositories · arXiv:2307.02435
-
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution 11 Jun 2023 · 1 repository · arXiv:2306.06755
-
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing 29 May 2023 · 0 repositories · arXiv:2305.18584
-
How Effective Are Neural Networks for Fixing Security Vulnerabilities 29 May 2023 · 1 repository · arXiv:2305.18607Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)
-
GrACE: Generation using Associated Code Edits 23 May 2023 · 0 repositories · arXiv:2305.14129
-
A Black-Box Attack on Code Models via Representation Nearest Neighbor Search 10 May 2023 · 0 repositories · arXiv:2305.05896
-
Enhancing Automated Program Repair through Fine-tuning and Prompt Engineering 16 Apr 2023 · 0 repositories · arXiv:2304.07840
-
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT 7 Mar 2023 · 1 repository · arXiv:2303.04226
Tasks archive 2025-07-28
20 shown of 41 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections