{"url":"/method/ceit","slug":"ceit","name":"CeiT","full_name":"Convolution-enhanced image Transformer","full_name_withheld":false,"description_markdown":"**Convolution-enhanced image Transformer** (**CeiT**) combines the advantages of CNNs in extracting low-level features, strengthening locality, and the advantages of Transformers in establishing long-range dependencies. Three modifications are made to the original Transformer: 1) instead of the straightforward tokenization from raw input images, we design an **Image-to-Tokens** (**I2T**) module that extracts patches from generated low-level features; 2) the feed-froward network in each encoder block is replaced with a **Locally-enhanced Feed-Forward** (**LeFF**) layer that promotes the correlation among neighbouring tokens in the spatial dimension; 3) a **Layer-wise Class token Attention** (**LCA**) is attached at the top of the Transformer that utilizes the multi-level representations.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2103.11816v2","title":"Incorporating Convolution Designs into Visual Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]}],"n_papers_tagged":2,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"DropKey for Vision Transformer","date":"2023-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/incorporating-convolution-designs-into-visual","title":"Incorporating Convolution Designs into Visual Transformers","date":"2021-03-22","arxiv_id":"2103.11816","n_code_links":3,"syntology":{"ran":9,"of":11,"unverified":2,"pointer_only":0}}],"papers_shown":2,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":2},{"task":"/task/human-object-interaction-detection","name":"Human-Object Interaction Detection","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1},{"task":"/task/object-detection-1","name":"object-detection","papers":1}],"tasks_shown":5,"n_tasks":5,"usage_by_year":[{"year":"2021","papers":1},{"year":"2023","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/ceit"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}