{"url":"/method/setr","slug":"setr","name":"SETR","full_name":"Segmentation Transformer","full_name_withheld":false,"description_markdown":"**Segmentation Transformer**, or **SETR**, is a [Transformer](https://paperswithcode.com/methods/category/transformers)-based segmentation model. The transformer-alone encoder treats an input image as a sequence of image patches represented by learned patch embedding, and transforms the sequence with global self-attention modeling for discriminative feature representation learning. Concretely, we first decompose an image into a grid of fixed-sized patches, forming a sequence of patches. With a linear embedding layer applied to the flattened pixel vectors of every patch, we then obtain a sequence of feature embedding vectors as the input to a transformer. Given the learned features from the encoder\r\ntransformer, a decoder is then used to recover the original image resolution. Crucially there is no downsampling in spatial resolution but global context modeling at every layer of the encoder transformer.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2012.15840v3","title":"Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Semantic Segmentation Models","url":"/methods/category/semantic-segmentation-models","pwc_aliases":["segmentation-models"]}],"n_papers_tagged":3,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"SAM for Poultry Science","date":"2023-05-17","arxiv_id":"2305.10254","n_code_links":0,"syntology":null},{"paper":null,"title":"Single Event Transition Risk: A Measure for Long Term Carbon Exposure","date":"2021-07-14","arxiv_id":"2107.06518","n_code_links":0,"syntology":null},{"paper":"/paper/rethinking-semantic-segmentation-from-a","title":"Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers","date":"2020-12-31","arxiv_id":"2012.15840","n_code_links":5,"syntology":null}],"papers_shown":3,"tasks":[{"task":"/task/segmentation","name":"Segmentation","papers":2},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":2},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/medical-image-segmentation","name":"Medical Image Segmentation","papers":1},{"task":"/task/object-tracking","name":"Object Tracking","papers":1},{"task":"/task/zero-shot-segmentation","name":"Zero Shot Segmentation","papers":1}],"tasks_shown":6,"n_tasks":6,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":1},{"year":"2023","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/setr"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}