{"url":"/method/musiq","slug":"musiq","name":"MUSIQ","full_name":"MUSIQ","full_name_withheld":false,"description_markdown":"**MUSIQ**, or **Multi-scale Image Quality Transformer**, is a [Transformer](https://paperswithcode.com/method/transformer)-based model for multi-scale image quality assessment. It processes native resolution images with varying sizes and aspect ratios. In MUSIQ, we construct a multi-scale image representation as input, including the native resolution image and its ARP resized variants.  Each image is split into fixed-size patches which are embedded by a patch encoding module (blue boxes). To capture 2D structure of the image and handle images of varying aspect ratios, the spatial embedding is encoded by hashing the patch position $(i,j)$ to $(t_{i},t_{j})$ within a grid of learnable embeddings (red boxes). Scale Embedding (green boxes) is introduced to capture scale information. The Transformer encoder takes the input tokens and performs multi-head self-attention. To predict the image quality, MUSIQ follows a common strategy in Transformers to add an [CLS] token to the sequence to represent the whole multi-scale input and the corresponding Transformer output is used as the final representation.","description_state":"present","introduced_year":null,"introduced_by":{"title":"MUSIQ: Multi-scale Image Quality Transformer","paper":"/paper/musiq-multi-scale-image-quality-transformer","first_author":"Junjie Ke","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/musiq-multi-scale-image-quality-transformer"},"source":{"url":"https://arxiv.org/abs/2108.05997v1","title":"MUSIQ: Multi-scale Image Quality Transformer","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Quality Models","url":"/methods/category/image-quality-models","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]}],"n_papers_tagged":4,"archive_num_papers":4,"papers_newest_first":[{"paper":null,"title":"Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling","date":"2025-05-17","arxiv_id":"2505.12048","n_code_links":0,"syntology":null},{"paper":null,"title":"NightHaze: Nighttime Image Dehazing via Self-Prior Learning","date":"2024-03-12","arxiv_id":"2403.07408","n_code_links":0,"syntology":null},{"paper":null,"title":"DIFFNAT: Improving Diffusion Image Quality Using Natural Image Statistics","date":"2023-11-16","arxiv_id":"2311.09753","n_code_links":0,"syntology":null},{"paper":"/paper/musiq-multi-scale-image-quality-transformer","title":"MUSIQ: Multi-scale Image Quality Transformer","date":"2021-08-12","arxiv_id":"2108.05997","n_code_links":2,"syntology":{"ran":5,"of":6,"unverified":1,"pointer_only":6}}],"papers_shown":4,"tasks":[{"task":"/task/super-resolution","name":"Super-Resolution","papers":2},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/image-dehazing","name":"Image Dehazing","papers":1},{"task":"/task/image-enhancement","name":"Image Enhancement","papers":1},{"task":"/task/image-generation","name":"Image Generation","papers":1},{"task":"/task/image-quality-assessment","name":"Image Quality Assessment","papers":1},{"task":"/task/image-super-resolution","name":"Image Super-Resolution","papers":1},{"task":"/task/unconditional-image-generation","name":"Unconditional Image Generation","papers":1},{"task":"/task/video-quality-assessment","name":"Video Quality Assessment","papers":1}],"tasks_shown":9,"n_tasks":9,"usage_by_year":[{"year":"2021","papers":1},{"year":"2023","papers":1},{"year":"2024","papers":1},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/musiq"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}