{"url":"/method/mvit","slug":"mvit","name":"MViT","full_name":"Multiscale Vision Transformer","full_name_withheld":false,"description_markdown":"**Multiscale Vision Transformer**, or **MViT**, is a [transformer](https://paperswithcode.com/method/transformer) architecture for modeling visual data such as images and videos. Unlike conventional transformers, which maintain a constant channel capacity and resolution throughout the network, Multiscale Transformers have several channel-resolution scale stages. Starting from the input resolution and a small channel dimension, the stages hierarchically expand the channel capacity while reducing the spatial resolution. This creates a multiscale pyramid of features with early layers operating at high spatial resolution to model simple low-level visual information, and deeper layers at spatially coarse, but complex, high-dimensional features.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Multiscale Vision Transformers","paper":"/paper/multiscale-vision-transformers","first_author":"Haoqi Fan","n_authors":7,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/multiscale-vision-transformers"},"source":{"url":"https://arxiv.org/abs/2104.11227v1","title":"Multiscale Vision Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]}],"n_papers_tagged":9,"archive_num_papers":9,"papers_newest_first":[{"paper":null,"title":"ROI-Aware Multiscale Cross-Attention Vision Transformer for Pest Image Identification","date":"2023-12-28","arxiv_id":"2312.16914","n_code_links":0,"syntology":null},{"paper":"/paper/pareprop-fast-parallelized-reversible","title":"PaReprop: Fast Parallelized Reversible Backpropagation","date":"2023-06-15","arxiv_id":"2306.09342","n_code_links":1,"syntology":null},{"paper":null,"title":"SVT: Supertoken Video Transformer for Efficient Video Understanding","date":"2023-04-01","arxiv_id":"2304.00325","n_code_links":0,"syntology":null},{"paper":null,"title":"Multi-Channel Vision Transformer for Epileptic Seizure Prediction","date":"2022-06-29","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Benchmarking Conventional Vision Models on Neuromorphic Fall Detection and Action Recognition Dataset","date":"2022-01-28","arxiv_id":"2201.12285","n_code_links":0,"syntology":null},{"paper":"/paper/improved-multiscale-vision-transformers-for","title":"MViTv2: Improved Multiscale Vision Transformers for Classification and Detection","date":"2021-12-02","arxiv_id":"2112.01526","n_code_links":9,"syntology":null},{"paper":"/paper/efficient-video-transformers-with-spatial","title":"Efficient Video Transformers with Spatial-Temporal Token Selection","date":"2021-11-23","arxiv_id":"2111.11591","n_code_links":1,"syntology":null},{"paper":"/paper/multi-modal-transformers-excel-at-class","title":"Class-agnostic Object Detection with Multi-modal Transformer","date":"2021-11-22","arxiv_id":"2111.11430","n_code_links":1,"syntology":null},{"paper":"/paper/multiscale-vision-transformers","title":"Multiscale Vision Transformers","date":"2021-04-22","arxiv_id":"2104.11227","n_code_links":8,"syntology":{"ran":13,"of":26,"unverified":13,"pointer_only":5}}],"papers_shown":9,"tasks":[{"task":"/task/action-recognition-in-videos","name":"Action Recognition","papers":3},{"task":"/task/video-recognition","name":"Video Recognition","papers":3},{"task":"/task/action-classification","name":"Action Classification","papers":2},{"task":"/task/benchmarking","name":"Benchmarking","papers":2},{"task":"/task/image-classification","name":"Image Classification","papers":2},{"task":"/task/object","name":"Object","papers":2},{"task":"/task/object-detection","name":"Object Detection","papers":2},{"task":"/task/class-agnostic-object-detection","name":"Class-agnostic Object Detection","papers":1},{"task":"/task/eeg-1","name":"EEG","papers":1},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":1},{"task":"/task/object-proposal-generation","name":"Object Proposal Generation","papers":1},{"task":"/task/open-world-object-detection","name":"Open World Object Detection","papers":1},{"task":"/task/prediction","name":"Prediction","papers":1},{"task":"/task/seizure-prediction","name":"Seizure prediction","papers":1},{"task":"/task/action-recognition","name":"Temporal Action Localization","papers":1},{"task":"/task/time-series-1","name":"Time Series","papers":1},{"task":"/task/video-classification","name":"Video Classification","papers":1},{"task":"/task/video-understanding","name":"Video Understanding","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1},{"task":"/task/object-detection-1","name":"object-detection","papers":1}],"tasks_shown":20,"n_tasks":20,"usage_by_year":[{"year":"2021","papers":4},{"year":"2022","papers":2},{"year":"2023","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/mvit"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}