Methods › Computer Vision › Vision Transformers › TNT
Transformer in Transformer
TNT
Introduced by Kai Han et al. in Transformer in Transformer
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Transformer is a type of self-attention-based neural networks originally applied for NLP tasks. Recently, pure transformer-based models are proposed to solve computer vision problems. These visual transformers usually view an image as a sequence of patches while they ignore the intrinsic structure information inside each patch. In this paper, we propose a novel Transformer-iN-Transformer (TNT) model for modeling both patch-level and pixel-level representation. In each TNT block, an outer transformer block is utilized to process patch embeddings, and an inner transformer block extracts local features from pixel embeddings. The pixel-level feature is projected to the space of patch embedding by a linear transformation layer and then added into the patch. By stacking the TNT blocks, we build the TNT model for image recognition.
Image source: Han et al.
Papers archive 2025-07-28
12 shown of 12, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Quadratic Gaussian Splatting for Efficient and Detailed Surface Reconstruction 25 Nov 2024 · 0 repositories · arXiv:2411.16392
-
Intensity Field Decomposition for Tissue-Guided Neural Tomography 1 Nov 2024 · 0 repositories · arXiv:2411.00900
-
Towards Minimal Targeted Updates of Language Models with Targeted Negative Training 19 Jun 2024 · 1 repository · arXiv:2406.13660
-
Swin transformers are robust to distribution and concept drift in endoscopy-based longitudinal rectal cancer assessment 6 May 2024 · 0 repositories · arXiv:2405.03762
-
Revolutionizing Traffic Sign Recognition: Unveiling the Potential of Vision Transformers 29 Apr 2024 · 0 repositories · arXiv:2404.19066
-
Nested-TNT: Hierarchical Vision Transformers with Multi-Scale Feature Processing 20 Apr 2024 · 0 repositories · arXiv:2404.13434
-
A Mini-Block Fisher Method for Deep Neural Networks 8 Feb 2022 · 0 repositories · arXiv:2202.04124
-
PyramidTNT: Improved Transformer-in-Transformer Baselines with Pyramid Architecture 4 Jan 2022 · 1 repository · arXiv:2201.00978
-
TnT Attacks! Universal Naturalistic Adversarial Patches Against Deep Neural Network Systems 19 Nov 2021 · 0 repositories · arXiv:2111.09999
-
Tree in Tree: from Decision Trees to Decision Graphs 1 Oct 2021 · 1 repository · arXiv:2110.00392
-
Tensor Normal Training for Deep Learning Models 5 Jun 2021 · 1 repository · arXiv:2106.02925Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)
-
Transformer in Transformer 27 Feb 2021 · 12 repositories · arXiv:2103.00112Syntology ran 16 of 24 samples · 8 unverified · 5 pointer-only (licence)
Tasks archive 2025-07-28
14 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections