Methods › Computer Vision › Image Model Blocks › Spatial Transformer
Spatial Transformer
Introduced by Max Jaderberg et al. in Spatial Transformer Networks
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
A Spatial Transformer is an image model block that explicitly allows the spatial manipulation of data within a convolutional neural network. It gives CNNs the ability to actively spatially transform feature maps, conditional on the feature map itself, without any extra training supervision or modification to the optimisation process. Unlike pooling layers, where the receptive fields are fixed and local, the spatial transformer module is a dynamic mechanism that can actively spatially transform an image (or a feature map) by producing an appropriate transformation for each input sample. The transformation is then performed on the entire feature map (non-locally) and can include scaling, cropping, rotations, as well as non-rigid deformations.
The architecture is shown in the Figure to the right. The input feature map U is passed to a localisation network which regresses the transformation parameters θ. The regular spatial grid G over V is transformed to the sampling grid T_θ(G), which is applied to U, producing the warped output feature map V. The combination of the localisation network and sampling mechanism defines a spatial transformer.
Papers archive 2025-07-28
30 shown of 169, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception 16 Jun 2025 · 0 repositories · arXiv:2506.13501
-
GuidedMorph: Two-Stage Deformable Registration for Breast MRI 19 May 2025 · 0 repositories · arXiv:2505.13414
-
EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition 14 Jan 2025 · 1 repository · arXiv:2501.08199
-
Neural encoding with affine feature response transforms 7 Jan 2025 · 1 repository · arXiv:2501.03741
-
A novel deep learning approach for facial emotion recognition: application to detecting emotional responses in elderly individuals with Alzheimer’s disease 30 Dec 2024 · 1 repository
-
Fixing the Perspective: A Critical Examination of Zero-1-to-3 24 Nov 2024 · 0 repositories · arXiv:2411.15706
-
ESC-MISR: Enhancing Spatial Correlations for Multi-Image Super-Resolution in Remote Sensing 7 Nov 2024 · 0 repositories · arXiv:2411.04706
-
Spatial Transformers for Radio Map Estimation 2 Nov 2024 · 0 repositories · arXiv:2411.01211
-
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction 24 Oct 2024 · 0 repositories · arXiv:2410.18962
-
Disambiguating Monocular Reconstruction of 3D Clothed Human with Spatial-Temporal Transformer 21 Oct 2024 · 0 repositories · arXiv:2410.16337
-
MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping 17 Sep 2024 · 1 repository · arXiv:2409.11316
-
Automatic facial axes standardization of 3D fetal ultrasound images 4 Sep 2024 · 0 repositories · arXiv:2409.02826
-
Improved 3D Whole Heart Geometry from Sparse CMR Slices 14 Aug 2024 · 1 repository · arXiv:2408.07532
-
Spatial Transformer Network YOLO Model for Agricultural Object Detection 31 Jul 2024 · 1 repository · arXiv:2407.21652
-
Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning 22 Jul 2024 · 0 repositories · arXiv:2407.15815
-
X-Recon: Learning-based Patient-specific High-Resolution CT Reconstruction from Orthogonal X-Ray Images 22 Jul 2024 · 1 repository · arXiv:2407.15356
-
Make Graph Neural Networks Great Again: A Generic Integration Paradigm of Topology-Free Patterns for Traffic Speed Prediction 24 Jun 2024 · 1 repository · arXiv:2406.16992
-
Infinite 3D Landmarks: Improving Continuous 2D Facial Landmark Detection 30 May 2024 · 0 repositories · arXiv:2405.20117
-
Vision-Language Modeling with Regularized Spatial Transformer Networks for All Weather Crosswind Landing of Aircraft 9 May 2024 · 0 repositories · arXiv:2405.05574
-
Efficient and Scalable Chinese Vector Font Generation via Component Composition 10 Apr 2024 · 0 repositories · arXiv:2404.06779
-
Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser 7 Mar 2024 · 1 repository · arXiv:2403.04444
-
The Paradox of Motion: Evidence for Spurious Correlations in Skeleton-based Gait Recognition Models 13 Feb 2024 · 0 repositories · arXiv:2402.08320
-
Simultaneous Alignment and Surface Regression Using Hybrid 2D-3D Networks for 3D Coherent Layer Segmentation of Retinal OCT Images with Full and Sparse Annotations 4 Dec 2023 · 1 repository · arXiv:2312.01726
-
MultiScale Spectral-Spatial Convolutional Transformer for Hyperspectral Image Classification 28 Oct 2023 · 0 repositories · arXiv:2310.18550
-
A Multi-Scale Spatial Transformer U-Net for Simultaneously Automatic Reorientation and Segmentation of 3D Nuclear Cardiac Images 16 Oct 2023 · 0 repositories · arXiv:2310.10095
-
Revisiting Data Augmentation for Rotational Invariance in Convolutional Neural Networks 12 Oct 2023 · 0 repositories · arXiv:2310.08429
-
UnitedHuman: Harnessing Multi-Source Data for High-Resolution Human Generation 25 Sep 2023 · 1 repository · arXiv:2309.14335
-
A Hierarchical Spatial Transformer for Massive Point Samples in Continuous Space 21 Sep 2023 · 1 repository
-
MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer 16 Sep 2023 · 1 repository · arXiv:2309.09067Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)
-
Test-Time Compensated Representation Learning for Extreme Traffic Forecasting 16 Sep 2023 · 0 repositories · arXiv:2309.09074
Tasks archive 2025-07-28
20 shown of 200 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections