{"url":"/method/affine-operator","slug":"affine-operator","name":"Affine Operator","full_name":"Affine Operator","full_name_withheld":false,"description_markdown":"The **Affine Operator** is an affine transformation layer introduced in the [ResMLP](https://paperswithcode.com/method/resmlp) architecture. This replaces [layer normalization](https://paperswithcode.com/method/layer-normalization), as in [Transformer based networks](https://paperswithcode.com/methods/category/transformers), which is possible since in the ResMLP, there are no [self-attention layers](https://paperswithcode.com/method/scaled) which makes training more stable - hence allowing a more simple affine transformation.\r\n\r\nThe affine operator is defined as:\r\n\r\n$$ \\operatorname{Aff}_{\\mathbf{\\alpha}, \\mathbf{\\beta}}(\\mathbf{x})=\\operatorname{Diag}(\\mathbf{\\alpha}) \\mathbf{x}+\\mathbf{\\beta} $$\r\n\r\nwhere $\\alpha$ and $\\beta$ are learnable weight vectors. This operation only rescales and shifts the input element-wise. This operation has several advantages over other normalization operations: first, as opposed to Layer Normalization, it has no cost at inference time, since it can absorbed in the adjacent linear layer. Second, as opposed to [BatchNorm](https://paperswithcode.com/method/batch-normalization) and Layer Normalization, the Aff operator does not depend on batch statistics.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2105.03404v2","title":"ResMLP: Feedforward networks for image classification with data-efficient training","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Feedforward Networks","url":"/methods/category/feedforward-networks","pwc_aliases":[]}],"n_papers_tagged":12,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/domain-influence-in-mri-medical-image","title":"Domain Influence in MRI Medical Image Segmentation: spatial versus k-space inputs","date":"2024-07-01","arxiv_id":"2407.01367","n_code_links":1,"syntology":null},{"paper":"/paper/sign-gradient-descent-based-neuronal-dynamics","title":"Sign Gradient Descent-based Neuronal Dynamics: ANN-to-SNN Conversion Beyond ReLU Network","date":"2024-07-01","arxiv_id":"2407.01645","n_code_links":1,"syntology":{"ran":12,"of":13,"unverified":1,"pointer_only":0}},{"paper":null,"title":"Compressing the Backward Pass of Large-Scale Neural Architectures by Structured Activation Pruning","date":"2023-11-28","arxiv_id":"2311.16883","n_code_links":0,"syntology":null},{"paper":null,"title":"Restore Translation Using Equivariant Neural Networks","date":"2023-06-29","arxiv_id":"2306.16938","n_code_links":0,"syntology":null},{"paper":"/paper/nilut-conditional-neural-implicit-3d-lookup","title":"NILUT: Conditional Neural Implicit 3D Lookup Tables for Image Enhancement","date":"2023-06-20","arxiv_id":"2306.11920","n_code_links":1,"syntology":null},{"paper":null,"title":"AutoInit: Automatic Initialization via Jacobian Tuning","date":"2022-06-27","arxiv_id":"2206.13568","n_code_links":0,"syntology":null},{"paper":null,"title":"Distributional Gaussian Processes Layers for Out-of-Distribution Detection","date":"2022-06-27","arxiv_id":"2206.13346","n_code_links":0,"syntology":null},{"paper":null,"title":"Boosting Adversarial Transferability of MLP-Mixer","date":"2022-04-26","arxiv_id":"2204.12204","n_code_links":0,"syntology":null},{"paper":"/paper/hire-mlp-vision-mlp-via-hierarchical","title":"Hire-MLP: Vision MLP via Hierarchical Rearrangement","date":"2021-08-30","arxiv_id":"2108.13341","n_code_links":10,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":1}},{"paper":"/paper/s-2-mlpv2-improved-spatial-shift-mlp","title":"S$^2$-MLPv2: Improved Spatial-Shift MLP Architecture for Vision","date":"2021-08-02","arxiv_id":"2108.01072","n_code_links":3,"syntology":{"ran":4,"of":4,"unverified":0,"pointer_only":2}},{"paper":"/paper/cyclemlp-a-mlp-like-architecture-for-dense","title":"CycleMLP: A MLP-like Architecture for Dense Prediction","date":"2021-07-21","arxiv_id":"2107.10224","n_code_links":8,"syntology":{"ran":8,"of":15,"unverified":7,"pointer_only":1}},{"paper":"/paper/resmlp-feedforward-networks-for-image","title":"ResMLP: Feedforward networks for image classification with data-efficient training","date":"2021-05-07","arxiv_id":"2105.03404","n_code_links":19,"syntology":{"ran":2,"of":7,"unverified":5,"pointer_only":0}}],"papers_shown":12,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":4},{"task":"/task/segmentation","name":"Segmentation","papers":3},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":3},{"task":"/task/image-classification","name":"image-classification","papers":3},{"task":"/task/object-detection","name":"Object Detection","papers":2},{"task":"/task/translation","name":"Translation","papers":2},{"task":"/task/object-detection-1","name":"object-detection","papers":2},{"task":"/task/adversarial-attack","name":"Adversarial Attack","papers":1},{"task":"/task/color-manipulation","name":"Color Manipulation","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/fine-grained-image-classification","name":"Fine-Grained Image Classification","papers":1},{"task":null,"name":"GPU","papers":1},{"task":"/task/gaussian-processes","name":"Gaussian Processes","papers":1},{"task":"/task/classification","name":"General Classification","papers":1},{"task":"/task/image-enhancement","name":"Image Enhancement","papers":1},{"task":"/task/image-segmentation","name":"Image Segmentation","papers":1},{"task":"/task/inductive-bias","name":"Inductive Bias","papers":1},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":1},{"task":"/task/machine-translation","name":"Machine Translation","papers":1},{"task":"/task/medical-image-segmentation","name":"Medical Image Segmentation","papers":1}],"tasks_shown":20,"n_tasks":26,"usage_by_year":[{"year":"2021","papers":4},{"year":"2022","papers":3},{"year":"2023","papers":3},{"year":"2024","papers":2}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/affine-operator"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}