Methods › Natural Language Processing › Transformers › CodeT5

CodeT5

32 papers tagged archive 2025-07-28

Introduced by Yue Wang et al. in CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

CodeT5 is a Transformer-based model for code understanding and generation based on the T5 architecture. It utilizes an identifier-aware pre-training objective that considers the crucial token type information (identifiers) from code. Specifically, the denoising Seq2Seq objective of T5 is extended with two identifier tagging and prediction tasks to enable the model to better leverage the token type information from programming languages, which are the identifiers assigned by developers. To improve the natural language-programming language alignment, a bimodal dual learning objective is used for a bidirectional conversion between natural language and programming language.

PaperSource

Papers archive 2025-07-28

30 shown of 32, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 41 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Code Generation8
Code Summarization6
Language Modelling5
Program Repair5
Vulnerability Detection4
Code Completion3
Code Translation3
Decoder3
Clone Detection2
Language Modeling2
Machine Translation2
Retrieval2
Small Language Model2
Translation2
Adversarial Attack1
Authorship Attribution1
Binary Classification1
Bug fixing1
Code Search1
Continual Learning1

Usage over time archive 2025-07-28

Papers per year tagged with CodeT5: 2015 to 2025, peak 15 15 0 2015: 1 paper 2015 2016: 0 papers 2016 2017: 0 papers 2017 2018: 0 papers 2018 2019: 0 papers 2019 2020: 0 papers 2020 2021: 1 paper 2021 2022: 0 papers 2022 2023: 15 papers 2023 2024: 7 papers 2024 2025: 8 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (32 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

TransformersAutoencoding TransformersCode Generation Transformers

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections