{"url":"/method/cait","slug":"cait","name":"CaiT","full_name":"Class-Attention in Image Transformers","full_name_withheld":false,"description_markdown":"**CaiT**, or **Class-Attention in Image Transformers**, is a type of [vision transformer](https://paperswithcode.com/methods/category/vision-transformer) with several design alterations upon the original [ViT](https://paperswithcode.com/method/vision-transformer). First a new layer scaling approach called [LayerScale](https://paperswithcode.com/method/layerscale) is used, adding a learnable diagonal matrix on output of each residual block, initialized close to (but not at) 0, which improves the training dynamics. Secondly, [class-attention layers](https://paperswithcode.com/method/ca) are introduced to the architecture. This creates an architecture where the transformer layers involving [self-attention](https://paperswithcode.com/method/scaled) between patches are explicitly separated from class-attention layers -- that are devoted to extract the content of the processed patches into a single vector so that it can be fed to a linear classifier.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2103.17239v2","title":"Going deeper with Image Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]}],"n_papers_tagged":5,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/spectralkd-understanding-and-optimizing","title":"SpectralKD: A Unified Framework for Interpreting and Distilling Vision Transformers via Spectral Analysis","date":"2024-12-26","arxiv_id":"2412.19055","n_code_links":1,"syntology":null},{"paper":null,"title":"Detecting Severity of Diabetic Retinopathy from Fundus Images: A Transformer Network-based Review","date":"2023-01-03","arxiv_id":"2301.00973","n_code_links":0,"syntology":null},{"paper":null,"title":"MaiT: Leverage Attention Masks for More Efficient Image Transformers","date":"2022-07-06","arxiv_id":"2207.03006","n_code_links":0,"syntology":null},{"paper":"/paper/mait-integrating-spatial-locality-into-image","title":"MaiT: integrating spatial locality into image transformers with attention masks","date":"2021-09-29","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/going-deeper-with-image-transformers","title":"Going deeper with Image Transformers","date":"2021-03-31","arxiv_id":"2103.17239","n_code_links":21,"syntology":{"ran":5,"of":11,"unverified":6,"pointer_only":2}}],"papers_shown":5,"tasks":[{"task":"/task/transfer-learning","name":"Transfer Learning","papers":2},{"task":"/task/image-classification","name":"Image Classification","papers":1},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1}],"tasks_shown":4,"n_tasks":4,"usage_by_year":[{"year":"2021","papers":2},{"year":"2022","papers":1},{"year":"2023","papers":1},{"year":"2024","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cait"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}