{"url":"/method/spatial-transformer","slug":"spatial-transformer","name":"Spatial Transformer","full_name":"Spatial Transformer","full_name_withheld":false,"description_markdown":"A **Spatial Transformer** is an image model block that explicitly allows the spatial manipulation of data within a [convolutional neural network](https://paperswithcode.com/methods/category/convolutional-neural-networks). It gives CNNs the ability to actively spatially transform feature maps, conditional on the feature map itself, without any extra training supervision or modification to the optimisation process. Unlike pooling layers, where the receptive fields are fixed and local, the spatial transformer module is a dynamic mechanism that can actively spatially transform an image (or a feature map) by producing an appropriate transformation for each input sample. The transformation is then performed on the entire feature map (non-locally) and can include scaling, cropping, rotations, as well as non-rigid deformations.\r\n\r\nThe architecture is shown in the Figure to the right. The input feature map $U$ is passed to a localisation network which regresses the transformation parameters $\\theta$. The regular spatial grid $G$ over $V$ is transformed to the sampling grid $T\\_{\\theta}\\left(G\\right)$, which is applied to $U$, producing the warped output feature map $V$. The combination of the localisation network and sampling mechanism defines a spatial transformer.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Spatial Transformer Networks","paper":"/paper/spatial-transformer-networks","first_author":"Max Jaderberg","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/spatial-transformer-networks"},"source":{"url":"http://arxiv.org/abs/1506.02025v3","title":"Spatial Transformer Networks","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/kevinzakka/spatial-transformer-network/blob/375f99046383316b18edfb5c575dc390c4ee3193/stn/transformer.py#L4","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Model Blocks","url":"/methods/category/image-model-blocks","pwc_aliases":[]}],"n_papers_tagged":169,"archive_num_papers":169,"papers_newest_first":[{"paper":null,"title":"FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception","date":"2025-06-16","arxiv_id":"2506.13501","n_code_links":0,"syntology":null},{"paper":null,"title":"GuidedMorph: Two-Stage Deformable Registration for Breast MRI","date":"2025-05-19","arxiv_id":"2505.13414","n_code_links":0,"syntology":null},{"paper":"/paper/emonext-an-adapted-convnext-for-facial-1","title":"EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition","date":"2025-01-14","arxiv_id":"2501.08199","n_code_links":1,"syntology":null},{"paper":"/paper/neural-encoding-with-affine-feature-response","title":"Neural encoding with affine feature response transforms","date":"2025-01-07","arxiv_id":"2501.03741","n_code_links":1,"syntology":null},{"paper":"/paper/a-novel-deep-learning-approach-for-facial","title":"A novel deep learning approach for facial emotion recognition: application to detecting emotional responses in elderly individuals with Alzheimer’s disease","date":"2024-12-30","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Fixing the Perspective: A Critical Examination of Zero-1-to-3","date":"2024-11-24","arxiv_id":"2411.15706","n_code_links":0,"syntology":null},{"paper":null,"title":"ESC-MISR: Enhancing Spatial Correlations for Multi-Image Super-Resolution in Remote Sensing","date":"2024-11-07","arxiv_id":"2411.04706","n_code_links":0,"syntology":null},{"paper":null,"title":"Spatial Transformers for Radio Map Estimation","date":"2024-11-02","arxiv_id":"2411.01211","n_code_links":0,"syntology":null},{"paper":null,"title":"Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction","date":"2024-10-24","arxiv_id":"2410.18962","n_code_links":0,"syntology":null},{"paper":null,"title":"Disambiguating Monocular Reconstruction of 3D Clothed Human with Spatial-Temporal Transformer","date":"2024-10-21","arxiv_id":"2410.16337","n_code_links":0,"syntology":null},{"paper":"/paper/msdnet-multi-scale-decoder-for-few-shot","title":"MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping","date":"2024-09-17","arxiv_id":"2409.11316","n_code_links":1,"syntology":null},{"paper":null,"title":"Automatic facial axes standardization of 3D fetal ultrasound images","date":"2024-09-04","arxiv_id":"2409.02826","n_code_links":0,"syntology":null},{"paper":"/paper/improved-3d-whole-heart-geometry-from-sparse","title":"Improved 3D Whole Heart Geometry from Sparse CMR Slices","date":"2024-08-14","arxiv_id":"2408.07532","n_code_links":1,"syntology":null},{"paper":"/paper/2407-21652","title":"Spatial Transformer Network YOLO Model for Agricultural Object Detection","date":"2024-07-31","arxiv_id":"2407.21652","n_code_links":1,"syntology":null},{"paper":null,"title":"Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning","date":"2024-07-22","arxiv_id":"2407.15815","n_code_links":0,"syntology":null},{"paper":"/paper/x-recon-learning-based-patient-specific-high","title":"X-Recon: Learning-based Patient-specific High-Resolution CT Reconstruction from Orthogonal X-Ray Images","date":"2024-07-22","arxiv_id":"2407.15356","n_code_links":1,"syntology":null},{"paper":"/paper/make-graph-neural-networks-great-again-a","title":"Make Graph Neural Networks Great Again: A Generic Integration Paradigm of Topology-Free Patterns for Traffic Speed Prediction","date":"2024-06-24","arxiv_id":"2406.16992","n_code_links":1,"syntology":null},{"paper":null,"title":"Infinite 3D Landmarks: Improving Continuous 2D Facial Landmark Detection","date":"2024-05-30","arxiv_id":"2405.20117","n_code_links":0,"syntology":null},{"paper":null,"title":"Vision-Language Modeling with Regularized Spatial Transformer Networks for All Weather Crosswind Landing of Aircraft","date":"2024-05-09","arxiv_id":"2405.05574","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficient and Scalable Chinese Vector Font Generation via Component Composition","date":"2024-04-10","arxiv_id":"2404.06779","n_code_links":0,"syntology":null},{"paper":"/paper/disentangled-diffusion-based-3d-human-pose","title":"Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser","date":"2024-03-07","arxiv_id":"2403.04444","n_code_links":1,"syntology":null},{"paper":null,"title":"The Paradox of Motion: Evidence for Spurious Correlations in Skeleton-based Gait Recognition Models","date":"2024-02-13","arxiv_id":"2402.08320","n_code_links":0,"syntology":null},{"paper":"/paper/simultaneous-alignment-and-surface-regression-1","title":"Simultaneous Alignment and Surface Regression Using Hybrid 2D-3D Networks for 3D Coherent Layer Segmentation of Retinal OCT Images with Full and Sparse Annotations","date":"2023-12-04","arxiv_id":"2312.01726","n_code_links":1,"syntology":null},{"paper":null,"title":"MultiScale Spectral-Spatial Convolutional Transformer for Hyperspectral Image Classification","date":"2023-10-28","arxiv_id":"2310.18550","n_code_links":0,"syntology":null},{"paper":null,"title":"A Multi-Scale Spatial Transformer U-Net for Simultaneously Automatic Reorientation and Segmentation of 3D Nuclear Cardiac Images","date":"2023-10-16","arxiv_id":"2310.10095","n_code_links":0,"syntology":null},{"paper":null,"title":"Revisiting Data Augmentation for Rotational Invariance in Convolutional Neural Networks","date":"2023-10-12","arxiv_id":"2310.08429","n_code_links":0,"syntology":null},{"paper":"/paper/unitedhuman-harnessing-multi-source-data-for","title":"UnitedHuman: Harnessing Multi-Source Data for High-Resolution Human Generation","date":"2023-09-25","arxiv_id":"2309.14335","n_code_links":1,"syntology":null},{"paper":"/paper/a-hierarchical-spatial-transformer-for","title":"A Hierarchical Spatial Transformer for Massive Point Samples in Continuous Space","date":"2023-09-21","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/mmst-vit-climate-change-aware-crop-yield","title":"MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer","date":"2023-09-16","arxiv_id":"2309.09067","n_code_links":1,"syntology":{"ran":5,"of":6,"unverified":1,"pointer_only":6}},{"paper":null,"title":"Test-Time Compensated Representation Learning for Extreme Traffic Forecasting","date":"2023-09-16","arxiv_id":"2309.09074","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/object","name":"Object","papers":12},{"task":"/task/object-detection","name":"Object Detection","papers":10},{"task":"/task/segmentation","name":"Segmentation","papers":10},{"task":"/task/object-detection-1","name":"object-detection","papers":10},{"task":"/task/decoder","name":"Decoder","papers":9},{"task":"/task/image-registration","name":"Image Registration","papers":9},{"task":"/task/image-classification","name":"Image Classification","papers":8},{"task":"/task/pose-estimation","name":"Pose Estimation","papers":8},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":8},{"task":"/task/image-classification","name":"image-classification","papers":8},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":7},{"task":"/task/classification","name":"General Classification","papers":7},{"task":"/task/representation-learning","name":"Representation Learning","papers":7},{"task":"/task/image-reconstruction","name":"Image Reconstruction","papers":5},{"task":"/task/super-resolution","name":"Super-Resolution","papers":5},{"task":"/task/classification-1","name":"Classification","papers":4},{"task":"/task/face-alignment","name":"Face Alignment","papers":4},{"task":"/task/person-re-identification","name":"Person Re-Identification","papers":4},{"task":"/task/scene-text-recognition","name":"Scene Text Recognition","papers":4},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":4}],"tasks_shown":20,"n_tasks":200,"usage_by_year":[{"year":"2015","papers":4},{"year":"2016","papers":7},{"year":"2017","papers":14},{"year":"2018","papers":13},{"year":"2019","papers":16},{"year":"2020","papers":23},{"year":"2021","papers":23},{"year":"2022","papers":25},{"year":"2023","papers":22},{"year":"2024","papers":18},{"year":"2025","papers":4}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/spatial-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}