{"url":"/method/mixer-layer","slug":"mixer-layer","name":"Mixer Layer","full_name":"MLP-Mixer Layer","full_name_withheld":false,"description_markdown":"A Mixer layer is a layer used in the MLP-Mixer architecture proposed by Tolstikhin et. al (2021) for computer vision. Mixer layers consist purely of MLPs, without convolutions or attention. It takes an input of embedded image patches (tokens), with its output having the same shape as its input, similar to that of a Vision Transformer encoder. As suggested by its name, Mixer layers \"mix\" tokens and channels through its \"token mixing\" and \"channel mixing\" MLPs contained the layer. It utilizes previous techniques by other architectures, such as layer normalization, skip-connections, and regularization methods.\r\n\r\nImage credit: Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., ... & Dosovitskiy, A. (2021). Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems, 34, 24261-24272.","description_state":"present","introduced_year":null,"introduced_by":{"title":"MLP-Mixer: An all-MLP Architecture for Vision","paper":"/paper/mlp-mixer-an-all-mlp-architecture-for-vision","first_author":"Ilya Tolstikhin","n_authors":12,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/mlp-mixer-an-all-mlp-architecture-for-vision"},"source":{"url":"https://arxiv.org/abs/2105.01601v4","title":"MLP-Mixer: An all-MLP Architecture for Vision","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/google-research/vision_transformer","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Model Blocks","url":"/methods/category/image-model-blocks","pwc_aliases":[]}],"n_papers_tagged":7,"archive_num_papers":7,"papers_newest_first":[{"paper":"/paper/triplemixer-a-3d-point-cloud-denoising-model","title":"TripleMixer: A 3D Point Cloud Denoising Model for Adverse Weather","date":"2024-08-25","arxiv_id":"2408.13802","n_code_links":1,"syntology":null},{"paper":"/paper/salsa-swift-adaptive-lightweight-self","title":"SALSA: Swift Adaptive Lightweight Self-Attention for Enhanced LiDAR Place Recognition","date":"2024-07-11","arxiv_id":"2407.08260","n_code_links":1,"syntology":null},{"paper":"/paper/ttms-fast-multi-level-tiny-time-mixers-for","title":"Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series","date":"2024-01-08","arxiv_id":"2401.03955","n_code_links":2,"syntology":{"ran":6,"of":6,"unverified":0,"pointer_only":0}},{"paper":"/paper/spatio-temporal-graph-mixformer-for-traffic","title":"Spatio-Temporal Graph Mixformer for Traffic Forecasting","date":"2023-10-15","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/a-mixer-layer-is-worth-one-graph-convolution","title":"Graph-Guided MLP-Mixer for Skeleton-Based Human Motion Prediction","date":"2023-04-07","arxiv_id":"2304.03532","n_code_links":0,"syntology":null},{"paper":null,"title":"MB-DECTNet: A Model-Based Unrolled Network for Accurate 3D DECT Reconstruction","date":"2023-02-01","arxiv_id":"2302.00577","n_code_links":0,"syntology":null},{"paper":"/paper/mlp-mixer-an-all-mlp-architecture-for-vision","title":"MLP-Mixer: An all-MLP Architecture for Vision","date":"2021-05-04","arxiv_id":"2105.01601","n_code_links":49,"syntology":{"ran":106,"of":134,"unverified":28,"pointer_only":36}}],"papers_shown":7,"tasks":[{"task":"/task/3d-place-recognition","name":"3D Place Recognition","papers":1},{"task":"/task/all","name":"All","papers":1},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":1},{"task":null,"name":"CPU","papers":1},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/few-shot-learning","name":"Few-Shot Learning","papers":1},{"task":"/task/human-pose-forecasting","name":"Human Pose Forecasting","papers":1},{"task":"/task/human-motion-prediction","name":"Human motion prediction","papers":1},{"task":"/task/image-classification","name":"Image Classification","papers":1},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":1},{"task":"/task/time-series-1","name":"Time Series","papers":1},{"task":"/task/time-series-forecasting","name":"Time Series Forecasting","papers":1},{"task":"/task/traffic-prediction","name":"Traffic Prediction","papers":1},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1},{"task":"/task/motion-prediction","name":"motion prediction","papers":1}],"tasks_shown":16,"n_tasks":16,"usage_by_year":[{"year":"2021","papers":1},{"year":"2023","papers":3},{"year":"2024","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/mixer-layer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}