{"url":"/method/tnt","slug":"tnt","name":"TNT","full_name":"Transformer in Transformer","full_name_withheld":false,"description_markdown":"[Transformer](https://paperswithcode.com/method/transformer) is a type of self-attention-based neural networks originally applied for NLP tasks. Recently, pure transformer-based models are proposed to solve computer vision problems. These visual transformers usually view an image as a sequence of patches while they ignore the intrinsic structure information inside each patch. In this paper, we propose a novel Transformer-iN-Transformer (TNT) model for modeling both patch-level and pixel-level representation. In each TNT block, an outer transformer block is utilized to process patch embeddings, and an inner transformer block extracts local features from pixel embeddings. The pixel-level feature is projected to the space of patch embedding by a linear transformation layer and then added into the patch. By stacking the TNT blocks, we build the TNT model for image recognition.\r\n\r\nImage source: [Han et al.](https://arxiv.org/pdf/2103.00112v1.pdf)","description_state":"present","introduced_year":null,"introduced_by":{"title":"Transformer in Transformer","paper":"/paper/transformer-in-transformer","first_author":"Kai Han","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/transformer-in-transformer"},"source":{"url":"https://arxiv.org/abs/2103.00112v3","title":"Transformer in Transformer","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Model Blocks","url":"/methods/category/image-model-blocks","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Models","url":"/methods/category/image-models","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Backbone Architectures","url":"/methods/category/backbone-architectures","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":12,"archive_num_papers":12,"papers_newest_first":[{"paper":null,"title":"Quadratic Gaussian Splatting for Efficient and Detailed Surface Reconstruction","date":"2024-11-25","arxiv_id":"2411.16392","n_code_links":0,"syntology":null},{"paper":null,"title":"Intensity Field Decomposition for Tissue-Guided Neural Tomography","date":"2024-11-01","arxiv_id":"2411.00900","n_code_links":0,"syntology":null},{"paper":"/paper/towards-minimal-targeted-updates-of-language","title":"Towards Minimal Targeted Updates of Language Models with Targeted Negative Training","date":"2024-06-19","arxiv_id":"2406.13660","n_code_links":1,"syntology":null},{"paper":null,"title":"Swin transformers are robust to distribution and concept drift in endoscopy-based longitudinal rectal cancer assessment","date":"2024-05-06","arxiv_id":"2405.03762","n_code_links":0,"syntology":null},{"paper":null,"title":"Revolutionizing Traffic Sign Recognition: Unveiling the Potential of Vision Transformers","date":"2024-04-29","arxiv_id":"2404.19066","n_code_links":0,"syntology":null},{"paper":null,"title":"Nested-TNT: Hierarchical Vision Transformers with Multi-Scale Feature Processing","date":"2024-04-20","arxiv_id":"2404.13434","n_code_links":0,"syntology":null},{"paper":null,"title":"A Mini-Block Fisher Method for Deep Neural Networks","date":"2022-02-08","arxiv_id":"2202.04124","n_code_links":0,"syntology":null},{"paper":"/paper/pyramidtnt-improved-transformer-in","title":"PyramidTNT: Improved Transformer-in-Transformer Baselines with Pyramid Architecture","date":"2022-01-04","arxiv_id":"2201.00978","n_code_links":1,"syntology":null},{"paper":null,"title":"TnT Attacks! Universal Naturalistic Adversarial Patches Against Deep Neural Network Systems","date":"2021-11-19","arxiv_id":"2111.09999","n_code_links":0,"syntology":null},{"paper":"/paper/tree-in-tree-from-decision-trees-to-decision","title":"Tree in Tree: from Decision Trees to Decision Graphs","date":"2021-10-01","arxiv_id":"2110.00392","n_code_links":1,"syntology":null},{"paper":"/paper/tensor-normal-training-for-deep-learning","title":"Tensor Normal Training for Deep Learning Models","date":"2021-06-05","arxiv_id":"2106.02925","n_code_links":1,"syntology":{"ran":2,"of":3,"unverified":1,"pointer_only":3}},{"paper":"/paper/transformer-in-transformer","title":"Transformer in Transformer","date":"2021-02-27","arxiv_id":"2103.00112","n_code_links":12,"syntology":{"ran":16,"of":24,"unverified":8,"pointer_only":5}}],"papers_shown":12,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":2},{"task":"/task/second-order-methods","name":"Second-order methods","papers":2},{"task":"/task/sentence","name":"Sentence","papers":2},{"task":"/task/3dgs","name":"3DGS","papers":1},{"task":"/task/autonomous-vehicles","name":"Autonomous Vehicles","papers":1},{"task":"/task/deep-learning","name":"Deep Learning","papers":1},{"task":"/task/evolutionary-algorithms","name":"Evolutionary Algorithms","papers":1},{"task":"/task/fine-grained-image-classification","name":"Fine-Grained Image Classification","papers":1},{"task":"/task/image-harmonization","name":"Image Harmonization","papers":1},{"task":"/task/nerf","name":"NeRF","papers":1},{"task":"/task/neural-rendering","name":"Neural Rendering","papers":1},{"task":"/task/surface-reconstruction","name":"Surface Reconstruction","papers":1},{"task":"/task/traffic-sign-recognition","name":"Traffic Sign Recognition","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1}],"tasks_shown":14,"n_tasks":14,"usage_by_year":[{"year":"2021","papers":4},{"year":"2022","papers":2},{"year":"2024","papers":6}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/tnt"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}