{"url":"/method/glu","slug":"glu","name":"Gated Linear Unit","full_name":"Gated Linear Unit","full_name_withheld":false,"description_markdown":"A Gated Linear Unit, or GLU computes:\r\n\r\n$$\r\n\\mathrm{GLU}(a, b) = a \\otimes \\sigma(b)\r\n$$\r\n\r\nIt is used in natural language processing architectures, for example the Gated CNN, because here $\\sigma(b)$ is the gate that control what information from $a$ is passed up to the following layer. Intuitively, for a language modeling task, the gating mechanism allows selection of words or features that are important for predicting the next word. The GLU also has non-linear capabilities, but has a linear path for the gradient so diminishes the vanishing gradient problem.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Language Modeling with Gated Convolutional Networks","paper":"/paper/language-modeling-with-gated-convolutional","first_author":"Yann N. Dauphin","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/language-modeling-with-gated-convolutional"},"source":{"url":"http://arxiv.org/abs/1612.08083v3","title":"Language Modeling with Gated Convolutional Networks","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Activation Functions","url":"/methods/category/activation-functions","pwc_aliases":[]}],"n_papers_tagged":798,"archive_num_papers":798,"papers_newest_first":[{"paper":null,"title":"LiLM-RDB-SFC: Lightweight Language Model with Relational Database-Guided DRL for Optimized SFC Provisioning","date":"2025-07-15","arxiv_id":"2507.10903","n_code_links":0,"syntology":null},{"paper":null,"title":"Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems","date":"2025-07-08","arxiv_id":"2507.05940","n_code_links":0,"syntology":null},{"paper":null,"title":"I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution","date":"2025-06-18","arxiv_id":"2506.17323","n_code_links":0,"syntology":null},{"paper":null,"title":"Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription","date":"2025-06-17","arxiv_id":"2506.14223","n_code_links":0,"syntology":null},{"paper":null,"title":"A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation","date":"2025-06-09","arxiv_id":"2506.08210","n_code_links":0,"syntology":null},{"paper":null,"title":"The Impact of Feature Scaling In Machine Learning: Effects on Regression and Classification Tasks","date":"2025-06-09","arxiv_id":"2506.08274","n_code_links":0,"syntology":null},{"paper":null,"title":"A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair","date":"2025-06-05","arxiv_id":"2506.04987","n_code_links":0,"syntology":null},{"paper":null,"title":"DuAL-Net: A Hybrid Framework for Alzheimer's Disease Prediction from Whole-Genome Sequencing via Local SNP Windows and Global Annotations","date":"2025-05-31","arxiv_id":"2506.00673","n_code_links":0,"syntology":null},{"paper":null,"title":"Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking","date":"2025-05-29","arxiv_id":"2505.23117","n_code_links":0,"syntology":null},{"paper":"/paper/shioenv-a-cli-behavior-capturing-environment","title":"ShIOEnv: A CLI Behavior-Capturing Environment Enabling Grammar-Guided Command Synthesis for Dataset Curation","date":"2025-05-23","arxiv_id":"2505.18374","n_code_links":1,"syntology":null},{"paper":null,"title":"Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification","date":"2025-05-22","arxiv_id":"2505.16338","n_code_links":0,"syntology":null},{"paper":"/paper/logicase-effective-test-case-generation-from","title":"LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming","date":"2025-05-21","arxiv_id":"2505.15039","n_code_links":0,"syntology":{"ran":11,"of":19,"unverified":8,"pointer_only":19}},{"paper":"/paper/eeg-to-text-translation-a-model-for","title":"EEG-to-Text Translation: A Model for Deciphering Human Brain Activity","date":"2025-05-20","arxiv_id":"2505.13936","n_code_links":1,"syntology":null},{"paper":"/paper/masking-in-multi-hop-qa-an-analysis-of-how","title":"Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation","date":"2025-05-16","arxiv_id":"2505.11754","n_code_links":1,"syntology":null},{"paper":null,"title":"Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits","date":"2025-05-14","arxiv_id":"2505.09407","n_code_links":0,"syntology":null},{"paper":null,"title":"Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization","date":"2025-05-08","arxiv_id":"2505.05070","n_code_links":0,"syntology":null},{"paper":"/paper/benchmarking-traditional-machine-learning-and","title":"Benchmarking Traditional Machine Learning and Deep Learning Models for Fault Detection in Power Transformers","date":"2025-05-07","arxiv_id":"2505.06295","n_code_links":1,"syntology":null},{"paper":"/paper/gascade-grouped-summarization-of-adverse-drug","title":"GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance","date":"2025-05-07","arxiv_id":"2505.04284","n_code_links":1,"syntology":null},{"paper":null,"title":"A review of DNA restriction-free overlapping sequence cloning techniques for synthetic biology","date":"2025-05-06","arxiv_id":"2505.03681","n_code_links":0,"syntology":null},{"paper":null,"title":"JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry","date":"2025-04-29","arxiv_id":"2504.20849","n_code_links":0,"syntology":null},{"paper":null,"title":"Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks","date":"2025-04-28","arxiv_id":"2504.19444","n_code_links":0,"syntology":null},{"paper":null,"title":"An Efficient Aerial Image Detection with Variable Receptive Fields","date":"2025-04-21","arxiv_id":"2504.15165","n_code_links":0,"syntology":null},{"paper":null,"title":"The Geometry of Self-Verification in a Task-Specific Reasoning Model","date":"2025-04-19","arxiv_id":"2504.14379","n_code_links":0,"syntology":null},{"paper":null,"title":"Can Moran Eigenvectors Improve Machine Learning of Spatial Data? Insights from Synthetic Data Validation","date":"2025-04-16","arxiv_id":"2504.12450","n_code_links":0,"syntology":null},{"paper":"/paper/enhancing-metabolic-syndrome-prediction-with","title":"Enhancing Metabolic Syndrome Prediction with Hybrid Data Balancing and Counterfactuals","date":"2025-04-09","arxiv_id":"2504.06987","n_code_links":1,"syntology":null},{"paper":"/paper/sigma-a-dataset-for-text-to-code-semantic","title":"Sigma: A dataset for text-to-code semantic parsing with statistical analysis","date":"2025-04-05","arxiv_id":"2504.04301","n_code_links":1,"syntology":null},{"paper":null,"title":"Advancing Sentiment Analysis in Tamil-English Code-Mixed Texts: Challenges and Transformer-Based Solutions","date":"2025-03-30","arxiv_id":"2503.23295","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Knowledge Graph Completion with Entity Neighborhood and Relation Context","date":"2025-03-29","arxiv_id":"2503.23205","n_code_links":0,"syntology":null},{"paper":"/paper/scaling-down-text-encoders-of-text-to-image","title":"Scaling Down Text Encoders of Text-to-Image Diffusion Models","date":"2025-03-25","arxiv_id":"2503.19897","n_code_links":1,"syntology":null},{"paper":"/paper/exploring-training-and-inference-scaling-laws","title":"Exploring Training and Inference Scaling Laws in Generative Retrieval","date":"2025-03-24","arxiv_id":"2503.18941","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":130},{"task":"/task/language-modeling","name":"Language Modeling","papers":101},{"task":"/task/question-answering","name":"Question Answering","papers":84},{"task":"/task/decoder","name":"Decoder","papers":79},{"task":"/task/text-generation","name":"Text Generation","papers":65},{"task":"/task/sentence","name":"Sentence","papers":57},{"task":"/task/translation","name":"Translation","papers":46},{"task":"/task/retrieval","name":"Retrieval","papers":42},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":42},{"task":"/task/machine-translation","name":"Machine Translation","papers":40},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":29},{"task":"/task/abstractive-text-summarization","name":"Abstractive Text Summarization","papers":24},{"task":"/task/semantic-parsing","name":"Semantic Parsing","papers":24},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":22},{"task":"/task/large-language-model","name":"Large Language Model","papers":20},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":20},{"task":"/task/code-generation","name":"Code Generation","papers":19},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":19},{"task":"/task/diversity","name":"Diversity","papers":19},{"task":"/task/text-summarization","name":"Text Summarization","papers":19}],"tasks_shown":20,"n_tasks":540,"usage_by_year":[{"year":"2016","papers":1},{"year":"2017","papers":1},{"year":"2018","papers":2},{"year":"2019","papers":12},{"year":"2020","papers":43},{"year":"2021","papers":120},{"year":"2022","papers":174},{"year":"2023","papers":215},{"year":"2024","papers":170},{"year":"2025","papers":60}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/glu"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}