{"url":"/method/layer-normalization","slug":"layer-normalization","name":"Layer Normalization","full_name":"Layer Normalization","full_name_withheld":false,"description_markdown":"Unlike [batch normalization](https://paperswithcode.com/method/batch-normalization), **Layer Normalization** directly estimates the normalization statistics from the summed inputs to the neurons within a hidden layer so the normalization does not introduce any new dependencies between training cases. It works well for [RNNs](https://paperswithcode.com/methods/category/recurrent-neural-networks) and improves both the training time and the generalization performance of several existing RNN models. More recently, it has been used with [Transformer](https://paperswithcode.com/methods/category/transformers) models.\r\n\r\nWe compute the layer normalization statistics over all the hidden units in the same layer as follows:\r\n\r\n$$ \\mu^{l} = \\frac{1}{H}\\sum^{H}\\_{i=1}a\\_{i}^{l} $$\r\n\r\n$$ \\sigma^{l} = \\sqrt{\\frac{1}{H}\\sum^{H}\\_{i=1}\\left(a\\_{i}^{l}-\\mu^{l}\\right)^{2}}  $$\r\n\r\nwhere $H$ denotes the number of hidden units in a layer. Under layer normalization, all the hidden units in a layer share the same normalization terms $\\mu$ and $\\sigma$, but different training cases have different normalization terms. Unlike batch normalization, layer normalization does not impose any constraint on the size of the mini-batch and it can be used in the pure online regime with batch size 1.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Layer Normalization","paper":"/paper/layer-normalization","first_author":"Jimmy Lei Ba","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/layer-normalization"},"source":{"url":"http://arxiv.org/abs/1607.06450v1","title":"Layer Normalization","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/CyberZHG/torch-layer-normalization/blob/89f405b60f53f85da6f03fe685c190ef394ce50c/torch_layer_normalization/layer_normalization.py#L8","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Normalization","url":"/methods/category/normalization","pwc_aliases":[]}],"n_papers_tagged":24980,"archive_num_papers":24985,"papers_newest_first":[{"paper":null,"title":"DASViT: Differentiable Architecture Search for Vision Transformer","date":"2025-07-17","arxiv_id":"2507.13079","n_code_links":0,"syntology":null},{"paper":"/paper/making-language-model-a-hierarchical","title":"Making Language Model a Hierarchical Classifier and Generator","date":"2025-07-17","arxiv_id":"2507.12930","n_code_links":1,"syntology":null},{"paper":"/paper/best-practices-for-large-scale-pixel-wise","title":"Best Practices for Large-Scale, Pixel-Wise Crop Mapping and Transfer Learning Workflows","date":"2025-07-16","arxiv_id":"2507.12590","n_code_links":1,"syntology":null},{"paper":"/paper/dvfl-net-a-lightweight-distilled-video-focal","title":"DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition","date":"2025-07-16","arxiv_id":"2507.12426","n_code_links":1,"syntology":null},{"paper":null,"title":"Biological Processing Units: Leveraging an Insect Connectome to Pioneer Biofidelic Neural Architectures","date":"2025-07-15","arxiv_id":"2507.10951","n_code_links":0,"syntology":null},{"paper":null,"title":"Generative Click-through Rate Prediction with Applications to Search Advertising","date":"2025-07-15","arxiv_id":"2507.11246","n_code_links":0,"syntology":null},{"paper":"/paper/hashed-watermark-as-a-filter-defeating","title":"Hashed Watermark as a Filter: Defeating Forging and Overwriting Attacks in Weight-based Neural Network Watermarking","date":"2025-07-15","arxiv_id":"2507.11137","n_code_links":1,"syntology":null},{"paper":"/paper/kv-latent-dimensional-level-kv-cache","title":"KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding","date":"2025-07-15","arxiv_id":"2507.11273","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":"/paper/langevin-flows-for-modeling-neural-latent","title":"Langevin Flows for Modeling Neural Latent Dynamics","date":"2025-07-15","arxiv_id":"2507.11531","n_code_links":1,"syntology":null},{"paper":null,"title":"Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI","date":"2025-07-13","arxiv_id":"2507.09702","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning from Synthetic Labs: Language Models as Auction Participants","date":"2025-07-12","arxiv_id":"2507.09083","n_code_links":0,"syntology":null},{"paper":null,"title":"Comparative Analysis of Vision Transformers and Traditional Deep Learning Approaches for Automated Pneumonia Detection in Chest X-Rays","date":"2025-07-11","arxiv_id":"2507.10589","n_code_links":0,"syntology":null},{"paper":null,"title":"A Wireless Foundation Model for Multi-Task Prediction","date":"2025-07-08","arxiv_id":"2507.05938","n_code_links":0,"syntology":null},{"paper":"/paper/agent-kb-leveraging-cross-domain-experience","title":"Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving","date":"2025-07-08","arxiv_id":"2507.06229","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems","date":"2025-07-08","arxiv_id":"2507.05940","n_code_links":0,"syntology":null},{"paper":null,"title":"Geo-Registration of Terrestrial LiDAR Point Clouds with Satellite Images without GNSS","date":"2025-07-08","arxiv_id":"2507.05999","n_code_links":0,"syntology":null},{"paper":"/paper/growing-transformers-modular-composition-and","title":"Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate","date":"2025-07-08","arxiv_id":"2507.07129","n_code_links":1,"syntology":null},{"paper":null,"title":"SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression","date":"2025-07-08","arxiv_id":"2507.05633","n_code_links":0,"syntology":null},{"paper":"/paper/tile-based-vit-inference-with-visual-cluster","title":"Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification","date":"2025-07-08","arxiv_id":"2507.06093","n_code_links":1,"syntology":null},{"paper":null,"title":"AI Generated Text Detection Using Instruction Fine-tuned Large Language and Transformer-Based Models","date":"2025-07-07","arxiv_id":"2507.05157","n_code_links":0,"syntology":null},{"paper":"/paper/emergent-semantics-beyond-token-embeddings","title":"Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations","date":"2025-07-07","arxiv_id":"2507.04886","n_code_links":1,"syntology":null},{"paper":null,"title":"Estimating Interventional Distributions with Uncertain Causal Graphs through Meta-Learning","date":"2025-07-07","arxiv_id":"2507.05526","n_code_links":0,"syntology":null},{"paper":"/paper/sv-drr-high-fidelity-novel-view-x-ray","title":"SV-DRR: High-Fidelity Novel View X-Ray Synthesis Using Diffusion Model","date":"2025-07-07","arxiv_id":"2507.05148","n_code_links":1,"syntology":null},{"paper":null,"title":"Behaviour Space Analysis of LLM-driven Meta-heuristic Discovery","date":"2025-07-04","arxiv_id":"2507.03605","n_code_links":0,"syntology":null},{"paper":null,"title":"CyberRAG: An agentic RAG cyber attack classification and reporting tool","date":"2025-07-03","arxiv_id":"2507.02424","n_code_links":0,"syntology":null},{"paper":"/paper/deepgesture-a-conversational-gesture","title":"DeepGesture: A conversational gesture synthesis system based on emotions and semantics","date":"2025-07-03","arxiv_id":"2507.03147","n_code_links":1,"syntology":null},{"paper":null,"title":"Fast and Simplex: 2-Simplicial Attention in Triton","date":"2025-07-03","arxiv_id":"2507.02754","n_code_links":0,"syntology":null},{"paper":null,"title":"Knowledge Protocol Engineering: A New Paradigm for AI in Domain-Specific Knowledge Work","date":"2025-07-03","arxiv_id":"2507.02760","n_code_links":0,"syntology":null},{"paper":null,"title":"AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation","date":"2025-07-02","arxiv_id":"2507.01961","n_code_links":0,"syntology":null},{"paper":"/paper/latent-chain-of-thought-decoding-the-depth","title":"Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer","date":"2025-07-02","arxiv_id":"2507.02199","n_code_links":1,"syntology":{"ran":0,"of":11,"unverified":11,"pointer_only":11}}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":3010},{"task":"/task/language-modeling","name":"Language Modeling","papers":2355},{"task":"/task/retrieval","name":"Retrieval","papers":1818},{"task":"/task/question-answering","name":"Question Answering","papers":1490},{"task":"/task/decoder","name":"Decoder","papers":1393},{"task":"/task/rag","name":"RAG","papers":1355},{"task":"/task/sentence","name":"Sentence","papers":1334},{"task":"/task/retrieval-augmented-generation","name":"Retrieval-augmented Generation","papers":1176},{"task":"/task/translation","name":"Translation","papers":1035},{"task":"/task/machine-translation","name":"Machine Translation","papers":925},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":862},{"task":"/task/large-language-model","name":"Large Language Model","papers":841},{"task":"/task/text-generation","name":"Text Generation","papers":763},{"task":"/task/image-classification","name":"Image Classification","papers":720},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":715},{"task":"/task/object-detection","name":"Object Detection","papers":668},{"task":"/task/representation-learning","name":"Representation Learning","papers":622},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":608},{"task":"/task/classification-1","name":"Classification","papers":597},{"task":"/task/object-detection-1","name":"object-detection","papers":597}],"tasks_shown":20,"n_tasks":2597,"usage_by_year":[{"year":"2016","papers":2},{"year":"2017","papers":26},{"year":"2018","papers":130},{"year":"2019","papers":1074},{"year":"2020","papers":2146},{"year":"2021","papers":3163},{"year":"2022","papers":3357},{"year":"2023","papers":5282},{"year":"2024","papers":6979},{"year":"2025","papers":2821}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/layer-normalization"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}