{"url":"/method/xcit-layer","slug":"xcit-layer","name":"XCiT Layer","full_name":"XCiT Layer","full_name_withheld":false,"description_markdown":"An **XCiT Layer** is the main building block of the [XCiT](https://paperswithcode.com/method/xcit) architecture which uses a [cross-covariance attention]() operator as its principal operation. The XCiT layer consists of three main blocks, each preceded by [LayerNorm](https://paperswithcode.com/method/layer-normalization) and followed by a [residual connection](https://paperswithcode.com/method/residual-connection): (i) the core [cross-covariance attention](https://paperswithcode.com/method/cross-covariance-attention) (XCA) operation, (ii) the [local patch interaction](https://paperswithcode.com/method/local-patch-interaction) (LPI) module, and (iii) a [feed-forward network](https://paperswithcode.com/method/feedforward-network) (FFN). By transposing the query-key interaction, the computational complexity of XCA is linear in the number of data elements N, rather than quadratic as in conventional self-attention.","description_state":"present","introduced_year":null,"introduced_by":{"title":"XCiT: Cross-Covariance Image Transformers","paper":"/paper/xcit-cross-covariance-image-transformers","first_author":"Alaaeldin El-Nouby","n_authors":11,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/xcit-cross-covariance-image-transformers"},"source":{"url":"https://arxiv.org/abs/2106.09681v2","title":"XCiT: Cross-Covariance Image Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Model Blocks","url":"/methods/category/image-model-blocks","pwc_aliases":[]}],"n_papers_tagged":4,"archive_num_papers":4,"papers_newest_first":[{"paper":null,"title":"CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision Models","date":"2022-10-14","arxiv_id":"2210.09223","n_code_links":0,"syntology":null},{"paper":"/paper/sima-simple-softmax-free-attention-for-vision","title":"SimA: Simple Softmax-free Attention for Vision Transformers","date":"2022-06-17","arxiv_id":"2206.08898","n_code_links":1,"syntology":{"ran":0,"of":3,"unverified":3,"pointer_only":0}},{"paper":"/paper/pe-former-pose-estimation-transformer","title":"PE-former: Pose Estimation Transformer","date":"2021-12-09","arxiv_id":"2112.04981","n_code_links":1,"syntology":null},{"paper":"/paper/xcit-cross-covariance-image-transformers","title":"XCiT: Cross-Covariance Image Transformers","date":"2021-06-17","arxiv_id":"2106.09681","n_code_links":12,"syntology":{"ran":3,"of":14,"unverified":11,"pointer_only":3}}],"papers_shown":4,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":3},{"task":"/task/image-classification","name":"image-classification","papers":2},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/pose-estimation","name":"Pose Estimation","papers":1},{"task":"/task/quantization","name":"Quantization","papers":1},{"task":"/task/self-supervised-image-classification","name":"Self-Supervised Image Classification","papers":1},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":1},{"task":"/task/object-detection-1","name":"object-detection","papers":1}],"tasks_shown":10,"n_tasks":10,"usage_by_year":[{"year":"2021","papers":2},{"year":"2022","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/xcit-layer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}