{"url":"/method/gpt-2","slug":"gpt-2","name":"GPT-2","full_name":"GPT-2","full_name_withheld":false,"description_markdown":"**GPT-2** is a [Transformer](https://paperswithcode.com/methods/category/transformers) architecture that was notable for its size (1.5 billion parameters) on its release. The model is pretrained on a [WebText dataset](https://paperswithcode.com/dataset/webtext) - text from 45 million website links. It largely follows the previous [GPT](https://paperswithcode.com/method/gpt) architecture with some modifications:\r\n\r\n- [Layer normalization](https://paperswithcode.com/method/layer-normalization) is moved to the input of each sub-block, similar to a\r\npre-activation residual network and an additional layer normalization was added after the final self-attention block. \r\n\r\n- A modified initialization which accounts for the accumulation on the residual path with model depth\r\nis used. Weights of residual layers are scaled at initialization by a factor of $1/\\sqrt{N}$ where $N$ is the number of residual layers. \r\n\r\n- The vocabulary is expanded to 50,257. The context size is expanded from 512 to 1024 tokens and\r\na larger batch size of 512 is used.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Language Models are Unsupervised Multitask Learners","paper":"/paper/language-models-are-unsupervised-multitask","first_author":"Alec Radford","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/language-models-are-unsupervised-multitask"},"source":{"url":"https://d4mucfpksywv.cloudfront.net/better-language-models/language-models.pdf","title":"Language Models are Unsupervised Multitask Learners","url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoregressive Transformers","url":"/methods/category/autoregressive-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":768,"archive_num_papers":768,"papers_newest_first":[{"paper":"/paper/the-automated-llm-speedrunning-benchmark","title":"The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements","date":"2025-06-27","arxiv_id":"2506.22419","n_code_links":1,"syntology":null},{"paper":null,"title":"M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models","date":"2025-06-17","arxiv_id":"2506.14532","n_code_links":0,"syntology":null},{"paper":"/paper/decomposing-mlp-activations-into","title":"Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization","date":"2025-06-12","arxiv_id":"2506.10920","n_code_links":1,"syntology":null},{"paper":null,"title":"A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning","date":"2025-06-11","arxiv_id":"2506.09429","n_code_links":0,"syntology":null},{"paper":null,"title":"Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models","date":"2025-06-08","arxiv_id":"2506.07121","n_code_links":0,"syntology":null},{"paper":null,"title":"Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective","date":"2025-06-05","arxiv_id":"2506.05166","n_code_links":0,"syntology":null},{"paper":null,"title":"An Exploratory Framework for Future SETI Applications: Detecting Generative Reactivity via Language Models","date":"2025-06-03","arxiv_id":"2506.02730","n_code_links":0,"syntology":null},{"paper":null,"title":"Rethinking the effects of data contamination in Code Intelligence","date":"2025-06-03","arxiv_id":"2506.02791","n_code_links":0,"syntology":null},{"paper":"/paper/how-neural-networks-organize-concepts","title":"How Neural Networks Organize Concepts: Introducing Concept Trajectory Analysis for Deep Learning Interpretability","date":"2025-06-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Power-of-Two (PoT) Weights in Large Language Models (LLMs)","date":"2025-05-31","arxiv_id":"2506.00315","n_code_links":0,"syntology":null},{"paper":null,"title":"Matryoshka Model Learning for Improved Elastic Student Models","date":"2025-05-29","arxiv_id":"2505.23337","n_code_links":0,"syntology":null},{"paper":null,"title":"Privacy-Preserving Chest X-ray Report Generation via Multimodal Federated Learning with ViT and GPT-2","date":"2025-05-27","arxiv_id":"2505.21715","n_code_links":0,"syntology":null},{"paper":null,"title":"Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents","date":"2025-05-26","arxiv_id":"2505.19494","n_code_links":0,"syntology":null},{"paper":null,"title":"Conversational Lexicography: Querying Lexicographic Data on Knowledge Graphs with SPARQL through Natural Language","date":"2025-05-26","arxiv_id":"2505.19971","n_code_links":0,"syntology":null},{"paper":null,"title":"ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining","date":"2025-05-26","arxiv_id":"2505.19893","n_code_links":0,"syntology":null},{"paper":null,"title":"Strong Membership Inference Attacks on Massive Datasets and (Moderately) Large Language Models","date":"2025-05-24","arxiv_id":"2505.18773","n_code_links":0,"syntology":null},{"paper":"/paper/adams-momentum-itself-can-be-a-normalizer-for","title":"AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training","date":"2025-05-22","arxiv_id":"2505.16363","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":0}},{"paper":null,"title":"The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm","date":"2025-05-22","arxiv_id":"2505.16932","n_code_links":0,"syntology":null},{"paper":"/paper/breaking-bad-tokens-detoxification-of-llms","title":"Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders","date":"2025-05-20","arxiv_id":"2505.14536","n_code_links":0,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":null,"title":"Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective","date":"2025-05-20","arxiv_id":"2505.14808","n_code_links":0,"syntology":null},{"paper":null,"title":"Scaling Laws for State Dynamics in Large Language Models","date":"2025-05-20","arxiv_id":"2505.14892","n_code_links":0,"syntology":null},{"paper":"/paper/vesselgpt-autoregressive-modeling-of-vascular","title":"VesselGPT: Autoregressive Modeling of Vascular Geometry","date":"2025-05-19","arxiv_id":"2505.13318","n_code_links":1,"syntology":null},{"paper":null,"title":"Can an Easy-to-Hard Curriculum Make Reasoning Emerge in Small Language Models? Evidence from a Four-Stage Curriculum on GPT-2","date":"2025-05-16","arxiv_id":"2505.11643","n_code_links":0,"syntology":null},{"paper":null,"title":"Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency","date":"2025-05-16","arxiv_id":"2505.13499","n_code_links":0,"syntology":null},{"paper":"/paper/are-sparse-autoencoders-useful-for-java","title":"Are Sparse Autoencoders Useful for Java Function Bug Detection?","date":"2025-05-15","arxiv_id":"2505.10375","n_code_links":1,"syntology":null},{"paper":null,"title":"Memorization-Compression Cycles Improve Generalization","date":"2025-05-13","arxiv_id":"2505.08727","n_code_links":0,"syntology":null},{"paper":"/paper/probability-consistency-in-large-language","title":"Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies","date":"2025-05-13","arxiv_id":"2505.08739","n_code_links":1,"syntology":null},{"paper":"/paper/llm-e-guess-can-llms-capabilities-advance","title":"LLM-e Guess: Can LLMs Capabilities Advance Without Hardware Progress?","date":"2025-05-07","arxiv_id":"2505.04075","n_code_links":1,"syntology":null},{"paper":"/paper/unidetox-universal-detoxification-of-large","title":"UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation","date":"2025-04-29","arxiv_id":"2504.20500","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":null,"title":"VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers","date":"2025-04-15","arxiv_id":"2504.11227","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":240},{"task":"/task/language-modeling","name":"Language Modeling","papers":183},{"task":"/task/text-generation","name":"Text Generation","papers":128},{"task":"/task/sentence","name":"Sentence","papers":49},{"task":"/task/question-answering","name":"Question Answering","papers":40},{"task":"/task/decoder","name":"Decoder","papers":39},{"task":"/task/large-language-model","name":"Large Language Model","papers":30},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":28},{"task":null,"name":"GPU","papers":26},{"task":"/task/retrieval","name":"Retrieval","papers":23},{"task":"/task/diversity","name":"Diversity","papers":20},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":20},{"task":"/task/in-context-learning","name":"In-Context Learning","papers":19},{"task":"/task/model","name":"model","papers":19},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":18},{"task":"/task/text-classification","name":"Text Classification","papers":17},{"task":"/task/translation","name":"Translation","papers":17},{"task":"/task/response-generation","name":"Response Generation","papers":16},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":16},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":15}],"tasks_shown":20,"n_tasks":398,"usage_by_year":[{"year":"2019","papers":31},{"year":"2020","papers":110},{"year":"2021","papers":119},{"year":"2022","papers":102},{"year":"2023","papers":138},{"year":"2024","papers":191},{"year":"2025","papers":77}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/gpt-2"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}