{"url":"/method/twins-pcpvt","slug":"twins-pcpvt","name":"Twins-PCPVT","full_name":"Twins-PCPVT","full_name_withheld":false,"description_markdown":"**Twins-PCPVT** is a type of [vision transformer](https://paperswithcode.com/methods/category/vision-transformer) that combines global attention, specifically the global sub-sampled attention as proposed in [Pyramid Vision Transformer](https://paperswithcode.com/method/pvt), with [conditional position encodings](https://paperswithcode.com/method/conditional-positional-encoding) (CPE) to replace the [absolute position encodings](https://paperswithcode.com/method/absolute-position-encodings) used in PVT.\r\n\r\nThe [position encoding generator](https://paperswithcode.com/method/positional-encoding-generator) (PEG), which generates the CPE, is placed after the first encoder block of each stage. The simplest form of PEG is used, i.e., a 2D [depth-wise convolution](https://paperswithcode.com/method/depthwise-convolution) without [batch normalization](https://paperswithcode.com/method/batch-normalization). For image-level classification, following [CPVT](https://paperswithcode.com/method/cpvt), the class token is removed and [global average pooling](https://paperswithcode.com/method/global-average-pooling) is used at the end of the stage. For other vision tasks, the design of PVT is followed.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2104.13840v4","title":"Twins: Revisiting the Design of Spatial Attention in Vision Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]}],"n_papers_tagged":1,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/twins-revisiting-spatial-attention-design-in","title":"Twins: Revisiting the Design of Spatial Attention in Vision Transformers","date":"2021-04-28","arxiv_id":"2104.13840","n_code_links":9,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":2}}],"papers_shown":1,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":1},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":1}],"tasks_shown":2,"n_tasks":2,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/twins-pcpvt"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}