{"url":"/method/mlp-mixer","slug":"mlp-mixer","name":"MLP-Mixer","full_name":"MLP-Mixer","full_name_withheld":false,"description_markdown":"The **MLP-Mixer** architecture (or “Mixer” for short) is an image architecture that doesn't use convolutions or self-attention. Instead, Mixer’s architecture is based entirely on multi-layer perceptrons (MLPs) that are repeatedly applied across either spatial locations or feature channels. Mixer relies only on basic matrix multiplication routines, changes to data layout (reshapes and transpositions), and scalar nonlinearities.\r\n\r\nIt accepts a sequence of linearly projected image patches (also referred to as tokens) shaped as a “patches × channels” table as an input, and maintains this dimensionality. Mixer makes use of two types of MLP layers: channel-mixing MLPs and token-mixing MLPs. The channel-mixing MLPs allow communication between different channels; they operate on each token independently and take individual rows of the table as inputs. The token-mixing MLPs allow communication between different spatial locations (tokens); they operate on each channel independently and take individual columns of the table as inputs. These two types of layers are interleaved to enable interaction of both input dimensions.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2105.01601v4","title":"MLP-Mixer: An all-MLP Architecture for Vision","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Models","url":"/methods/category/image-models","pwc_aliases":[]}],"n_papers_tagged":96,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Knowledge Distillation for Reservoir-based Classifier: Human Activity Recognition","date":"2025-05-29","arxiv_id":"2505.22985","n_code_links":0,"syntology":null},{"paper":null,"title":"Structured Initialization for Vision Transformers","date":"2025-05-26","arxiv_id":"2505.19985","n_code_links":0,"syntology":null},{"paper":null,"title":"Unified Cross-Modal Attention-Mixer Based Structural-Functional Connectomics Fusion for Neuropsychiatric Disorder Diagnosis","date":"2025-05-21","arxiv_id":"2505.15139","n_code_links":0,"syntology":null},{"paper":"/paper/med3dvlm-an-efficient-vision-language-model","title":"Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis","date":"2025-03-25","arxiv_id":"2503.20047","n_code_links":1,"syntology":{"ran":2,"of":12,"unverified":10,"pointer_only":0}},{"paper":"/paper/lightweight-models-for-emotional-analysis-in","title":"Lightweight Models for Emotional Analysis in Video","date":"2025-03-13","arxiv_id":"2503.10530","n_code_links":1,"syntology":null},{"paper":null,"title":"KAN-Mixers: a new deep learning architecture for image classification","date":"2025-03-11","arxiv_id":"2503.08939","n_code_links":0,"syntology":null},{"paper":null,"title":"Fast Jet Tagging with MLP-Mixers on FPGAs","date":"2025-03-05","arxiv_id":"2503.03103","n_code_links":0,"syntology":null},{"paper":"/paper/temporal-graph-mlp-mixer-for-spatio-temporal","title":"Temporal Graph MLP Mixer for Spatio-Temporal Forecasting","date":"2025-01-17","arxiv_id":"2501.10214","n_code_links":1,"syntology":null},{"paper":null,"title":"Towards Ideal Temporal Graph Neural Networks: Evaluations and Conclusions after 10,000 GPU Hours","date":"2024-12-28","arxiv_id":"2412.20256","n_code_links":0,"syntology":null},{"paper":"/paper/wpmixer-efficient-multi-resolution-mixing-for","title":"WPMixer: Efficient Multi-Resolution Mixing for Long-Term Time Series Forecasting","date":"2024-12-22","arxiv_id":"2412.17176","n_code_links":1,"syntology":{"ran":3,"of":5,"unverified":2,"pointer_only":0}},{"paper":"/paper/libragrad-balancing-gradient-flow-for","title":"LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions","date":"2024-11-24","arxiv_id":"2411.16760","n_code_links":1,"syntology":null},{"paper":null,"title":"LoRA-FAIR: Federated LoRA Fine-Tuning with Aggregation and Initialization Refinement","date":"2024-11-22","arxiv_id":"2411.14961","n_code_links":0,"syntology":null},{"paper":"/paper/improving-3d-medical-image-segmentation-at","title":"Improving 3D Medical Image Segmentation at Boundary Regions using Local Self-attention and Global Volume Mixing","date":"2024-10-20","arxiv_id":"2410.15360","n_code_links":1,"syntology":null},{"paper":null,"title":"Interpolated-MLPs: Controllable Inductive Bias","date":"2024-10-12","arxiv_id":"2410.09655","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Foundation Models for the Industrial Forecasting of Chemical Kinetics","date":"2024-08-20","arxiv_id":"2408.10720","n_code_links":0,"syntology":null},{"paper":"/paper/benchmarking-tree-species-classification-from","title":"Benchmarking tree species classification from proximally-sensed laser scanning data: introducing the FOR-species20K dataset","date":"2024-08-12","arxiv_id":"2408.06507","n_code_links":2,"syntology":null},{"paper":"/paper/hyperaggregation-aggregating-over-graph-edges","title":"HyperAggregation: Aggregating over Graph Edges with Hypernetworks","date":"2024-07-16","arxiv_id":"2407.11596","n_code_links":1,"syntology":null},{"paper":"/paper/sign-gradient-descent-based-neuronal-dynamics","title":"Sign Gradient Descent-based Neuronal Dynamics: ANN-to-SNN Conversion Beyond ReLU Network","date":"2024-07-01","arxiv_id":"2407.01645","n_code_links":1,"syntology":{"ran":12,"of":13,"unverified":1,"pointer_only":0}},{"paper":"/paper/demonstrating-the-efficacy-of-kolmogorov","title":"Demonstrating the Efficacy of Kolmogorov-Arnold Networks in Vision Tasks","date":"2024-06-21","arxiv_id":"2406.14916","n_code_links":1,"syntology":null},{"paper":"/paper/hierarchical-associative-memory-parallelized","title":"Hierarchical Associative Memory, Parallelized MLP-Mixer, and Symmetry Breaking","date":"2024-06-18","arxiv_id":"2406.12220","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/exploring-the-enigma-of-neural-dynamics","title":"Exploring the Enigma of Neural Dynamics Through A Scattering-Transform Mixer Landscape for Riemannian Manifold","date":"2024-05-25","arxiv_id":"2405.16357","n_code_links":1,"syntology":{"ran":11,"of":14,"unverified":3,"pointer_only":14}},{"paper":"/paper/mlps-learn-in-context","title":"MLPs Learn In-Context on Regression and Classification Tasks","date":"2024-05-24","arxiv_id":"2405.15618","n_code_links":2,"syntology":{"ran":2,"of":16,"unverified":14,"pointer_only":0}},{"paper":null,"title":"Transformer-based Federated Learning for Multi-Label Remote Sensing Image Classification","date":"2024-05-24","arxiv_id":"2405.15405","n_code_links":0,"syntology":null},{"paper":null,"title":"Node Centrality Approximation For Large Networks Based On Inductive Graph Neural Networks","date":"2024-03-08","arxiv_id":"2403.04977","n_code_links":0,"syntology":null},{"paper":"/paper/ninformer-a-network-in-network-transformer","title":"NiNformer: A Network in Network Transformer with Token Mixing Generated Gating Function","date":"2024-03-04","arxiv_id":"2403.02411","n_code_links":1,"syntology":null},{"paper":"/paper/mixer-is-more-than-just-a-model","title":"Mixer is more than just a model","date":"2024-02-28","arxiv_id":"2402.18007","n_code_links":0,"syntology":null},{"paper":"/paper/packd-pattern-clustered-knowledge","title":"PaCKD: Pattern-Clustered Knowledge Distillation for Compressing Memory Access Prediction Models","date":"2024-02-21","arxiv_id":"2402.13441","n_code_links":1,"syntology":null},{"paper":"/paper/multilinear-mixture-of-experts-scalable","title":"Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization","date":"2024-02-19","arxiv_id":"2402.12550","n_code_links":2,"syntology":{"ran":4,"of":4,"unverified":0,"pointer_only":4}},{"paper":"/paper/m2-mixer-a-multimodal-mixer-with-multi-head","title":"M2-Mixer: A Multimodal Mixer with Multi-head Loss for Classification from Multimodal Data","date":"2024-01-22","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/patchad-patch-based-mlp-mixer-for-time-series","title":"PatchAD: A Lightweight Patch-based MLP-Mixer for Time Series Anomaly Detection","date":"2024-01-18","arxiv_id":"2401.09793","n_code_links":1,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":0}}],"papers_shown":30,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":23},{"task":"/task/image-classification","name":"image-classification","papers":16},{"task":"/task/inductive-bias","name":"Inductive Bias","papers":9},{"task":"/task/object-detection","name":"Object Detection","papers":7},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":7},{"task":"/task/time-series-1","name":"Time Series","papers":6},{"task":"/task/object-detection-1","name":"object-detection","papers":6},{"task":"/task/classification-1","name":"Classification","papers":5},{"task":"/task/medical-image-analysis","name":"Medical Image Analysis","papers":4},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":4},{"task":"/task/all","name":"All","papers":3},{"task":"/task/benchmarking","name":"Benchmarking","papers":3},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":3},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":3},{"task":null,"name":"GPU","papers":3},{"task":"/task/image-to-image-translation","name":"Image-to-Image Translation","papers":3},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3},{"task":"/task/language-modelling","name":"Language Modelling","papers":3},{"task":"/task/time-series-forecasting","name":"Time Series Forecasting","papers":3}],"tasks_shown":20,"n_tasks":138,"usage_by_year":[{"year":"2021","papers":26},{"year":"2022","papers":21},{"year":"2023","papers":17},{"year":"2024","papers":24},{"year":"2025","papers":8}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/mlp-mixer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}