Methods › General › Normalization › Layer Normalization
Layer Normalization
Introduced by Jimmy Lei Ba et al. in Layer Normalization
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs to the neurons within a hidden layer so the normalization does not introduce any new dependencies between training cases. It works well for RNNs and improves both the training time and the generalization performance of several existing RNN models. More recently, it has been used with Transformer models.
We compute the layer normalization statistics over all the hidden units in the same layer as follows:
μˡ = 1/H∑ᴴᵢ₌₁aᵢˡ
σˡ = √(1/H∑ᴴᵢ₌₁(aᵢˡ-μˡ)²)
where H denotes the number of hidden units in a layer. Under layer normalization, all the hidden units in a layer share the same normalization terms μ and σ, but different training cases have different normalization terms. Unlike batch normalization, layer normalization does not impose any constraint on the size of the mini-batch and it can be used in the pure online regime with batch size 1.
Papers archive 2025-07-28
30 shown of 24,980, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
DASViT: Differentiable Architecture Search for Vision Transformer 17 Jul 2025 · 0 repositories · arXiv:2507.13079
-
Making Language Model a Hierarchical Classifier and Generator 17 Jul 2025 · 1 repository · arXiv:2507.12930
-
Best Practices for Large-Scale, Pixel-Wise Crop Mapping and Transfer Learning Workflows 16 Jul 2025 · 1 repository · arXiv:2507.12590
-
DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition 16 Jul 2025 · 1 repository · arXiv:2507.12426
-
Biological Processing Units: Leveraging an Insect Connectome to Pioneer Biofidelic Neural Architectures 15 Jul 2025 · 0 repositories · arXiv:2507.10951
-
Generative Click-through Rate Prediction with Applications to Search Advertising 15 Jul 2025 · 0 repositories · arXiv:2507.11246
-
Hashed Watermark as a Filter: Defeating Forging and Overwriting Attacks in Weight-based Neural Network Watermarking 15 Jul 2025 · 1 repository · arXiv:2507.11137
-
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding 15 Jul 2025 · 1 repository · arXiv:2507.11273Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
Langevin Flows for Modeling Neural Latent Dynamics 15 Jul 2025 · 1 repository · arXiv:2507.11531
-
Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI 13 Jul 2025 · 0 repositories · arXiv:2507.09702
-
Learning from Synthetic Labs: Language Models as Auction Participants 12 Jul 2025 · 0 repositories · arXiv:2507.09083
-
Comparative Analysis of Vision Transformers and Traditional Deep Learning Approaches for Automated Pneumonia Detection in Chest X-Rays 11 Jul 2025 · 0 repositories · arXiv:2507.10589
-
A Wireless Foundation Model for Multi-Task Prediction 8 Jul 2025 · 0 repositories · arXiv:2507.05938
-
Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving 8 Jul 2025 · 1 repository · arXiv:2507.06229Syntology ran 1 of 1 samples · 0 unverified
-
Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems 8 Jul 2025 · 0 repositories · arXiv:2507.05940
-
Geo-Registration of Terrestrial LiDAR Point Clouds with Satellite Images without GNSS 8 Jul 2025 · 0 repositories · arXiv:2507.05999
-
Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate 8 Jul 2025 · 1 repository · arXiv:2507.07129
-
SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression 8 Jul 2025 · 0 repositories · arXiv:2507.05633
-
Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification 8 Jul 2025 · 1 repository · arXiv:2507.06093
-
AI Generated Text Detection Using Instruction Fine-tuned Large Language and Transformer-Based Models 7 Jul 2025 · 0 repositories · arXiv:2507.05157
-
Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations 7 Jul 2025 · 1 repository · arXiv:2507.04886
-
Estimating Interventional Distributions with Uncertain Causal Graphs through Meta-Learning 7 Jul 2025 · 0 repositories · arXiv:2507.05526
-
SV-DRR: High-Fidelity Novel View X-Ray Synthesis Using Diffusion Model 7 Jul 2025 · 1 repository · arXiv:2507.05148
-
Behaviour Space Analysis of LLM-driven Meta-heuristic Discovery 4 Jul 2025 · 0 repositories · arXiv:2507.03605
-
CyberRAG: An agentic RAG cyber attack classification and reporting tool 3 Jul 2025 · 0 repositories · arXiv:2507.02424
-
DeepGesture: A conversational gesture synthesis system based on emotions and semantics 3 Jul 2025 · 1 repository · arXiv:2507.03147
-
Fast and Simplex: 2-Simplicial Attention in Triton 3 Jul 2025 · 0 repositories · arXiv:2507.02754
-
Knowledge Protocol Engineering: A New Paradigm for AI in Domain-Specific Knowledge Work 3 Jul 2025 · 0 repositories · arXiv:2507.02760
-
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation 2 Jul 2025 · 0 repositories · arXiv:2507.01961
-
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer 2 Jul 2025 · 1 repository · arXiv:2507.02199Syntology ran 0 of 11 samples · 11 unverified · 11 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 2,597 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Language Modelling | 3,010 |
| Language Modeling | 2,355 |
| Retrieval | 1,818 |
| Question Answering | 1,490 |
| Decoder | 1,393 |
| RAG | 1,355 |
| Sentence | 1,334 |
| Retrieval-augmented Generation | 1,176 |
| Translation | 1,035 |
| Machine Translation | 925 |
| Semantic Segmentation | 862 |
| Large Language Model | 841 |
| Text Generation | 763 |
| Image Classification | 720 |
| Transfer Learning | 715 |
| Object Detection | 668 |
| Representation Learning | 622 |
| Sentiment Analysis | 608 |
| Classification | 597 |
| object-detection | 597 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections