{"url":"/method/na","slug":"na","name":"Neighborhood Attention","full_name":"Neighborhood Attention","full_name_withheld":false,"description_markdown":"Neighborhood Attention is a restricted self attention pattern in which each token's receptive field is limited to its nearest neighboring pixels. It was proposed in [Neighborhood Attention Transformer](https://paperswithcode.com/paper/neighborhood-attention-transformer) as an alternative to other local attention mechanisms used in Hierarchical Vision Transformers.\r\n\r\nNA is in concept similar to [stand alone self attention (SASA)](https://paperswithcode.com/method/sasa), in that both can be implemented with a raster scan sliding window operation over the key value pair. However, NA would require a modification to handle corner pixels, which helps maintain a fixed receptive field size and an increased number of relative positions.\r\n\r\nThe primary challenge in experimenting with both NA and SASA has been computation. Simply extracting key values for each query is slow, takes up a large amount of memory, and is eventually intractable at scale. NA was therefore implemented through a new CUDA extension to PyTorch, [NATTEN](https://github.com/SHI-Labs/NATTEN).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Neighborhood Attention Transformer","paper":"/paper/neighborhood-attention-transformer","first_author":"Ali Hassani","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/neighborhood-attention-transformer"},"source":{"url":"https://arxiv.org/abs/2204.07143v5","title":"Neighborhood Attention Transformer","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Attention Mechanisms","url":"/methods/category/attention-mechanisms","pwc_aliases":["attention-mechanisms-1"]},{"area":"General","area_id":"general","collection":"Attention Modules","url":"/methods/category/attention-modules","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Attention Patterns","url":"/methods/category/attention-patterns","pwc_aliases":["factorized-attention"]}],"n_papers_tagged":21,"archive_num_papers":21,"papers_newest_first":[{"paper":"/paper/2505-11157","title":"Attention on the Sphere","date":"2025-05-16","arxiv_id":"2505.11157","n_code_links":1,"syntology":null},{"paper":"/paper/generalized-neighborhood-attention-multi","title":"Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light","date":"2025-04-23","arxiv_id":"2504.16922","n_code_links":1,"syntology":null},{"paper":"/paper/medical-image-classification-with-kan","title":"Medical Image Classification with KAN-Integrated Transformers and Dilated Neighborhood Attention","date":"2025-02-19","arxiv_id":"2502.13693","n_code_links":1,"syntology":null},{"paper":"/paper/d3rm-a-discrete-denoising-diffusion","title":"D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription","date":"2025-01-09","arxiv_id":"2501.05068","n_code_links":1,"syntology":null},{"paper":null,"title":"PIGUIQA: A Physical Imaging Guided Perceptual Framework for Underwater Image Quality Assessment","date":"2024-12-20","arxiv_id":"2412.15527","n_code_links":0,"syntology":null},{"paper":"/paper/edgenat-transformer-for-efficient-edge","title":"EdgeNAT: Transformer for Efficient Edge Detection","date":"2024-08-20","arxiv_id":"2408.10527","n_code_links":1,"syntology":null},{"paper":null,"title":"Compression-Realized Deep Structural Network for Video Quality Enhancement","date":"2024-05-10","arxiv_id":"2405.06342","n_code_links":0,"syntology":null},{"paper":null,"title":"A Point-Based Approach to Efficient LiDAR Multi-Task Perception","date":"2024-04-19","arxiv_id":"2404.12798","n_code_links":0,"syntology":null},{"paper":"/paper/deblurdinat-a-lightweight-and-effective","title":"DeblurDiNAT: A Compact Model with Exceptional Generalization and Visual Fidelity on Unseen Domains","date":"2024-03-19","arxiv_id":"2403.13163","n_code_links":1,"syntology":null},{"paper":"/paper/faster-neighborhood-attention-reducing-the-o","title":"Faster Neighborhood Attention: Reducing the O(n^2) Cost of Self Attention at the Threadblock Level","date":"2024-03-07","arxiv_id":"2403.04690","n_code_links":1,"syntology":{"ran":1,"of":3,"unverified":2,"pointer_only":0}},{"paper":"/paper/nac-tcn-temporal-convolutional-networks-with","title":"NAC-TCN: Temporal Convolutional Networks with Causal Dilated Neighborhood Attention for Emotion Understanding","date":"2023-12-12","arxiv_id":"2312.07507","n_code_links":1,"syntology":null},{"paper":"/paper/sdlformer-a-sparse-and-dense-locality","title":"SDLFormer: A Sparse and Dense Locality-enhanced Transformer for Accelerated MR Image Reconstruction","date":"2023-08-08","arxiv_id":"2308.04262","n_code_links":1,"syntology":null},{"paper":"/paper/modet-learning-deformable-image-registration","title":"ModeT: Learning Deformable Image Registration via Motion Decomposition Transformer","date":"2023-06-09","arxiv_id":"2306.05688","n_code_links":1,"syntology":null},{"paper":"/paper/dilated-unet-a-fast-and-accurate-medical","title":"Dilated-UNet: A Fast and Accurate Medical Image Segmentation Approach using a Dilated Transformer and U-Net Architecture","date":"2023-04-22","arxiv_id":"2304.11450","n_code_links":1,"syntology":null},{"paper":"/paper/incorporating-transformer-designs-into","title":"Incorporating Transformer Designs into Convolutions for Lightweight Image Super-Resolution","date":"2023-03-25","arxiv_id":"2303.14324","n_code_links":1,"syntology":null},{"paper":"/paper/oneformer-one-transformer-to-rule-universal","title":"OneFormer: One Transformer to Rule Universal Image Segmentation","date":"2022-11-10","arxiv_id":"2211.06220","n_code_links":4,"syntology":{"ran":0,"of":5,"unverified":5,"pointer_only":0}},{"paper":"/paper/stylenat-giving-each-head-a-new-perspective","title":"StyleNAT: Giving Each Head a New Perspective","date":"2022-11-10","arxiv_id":"2211.05770","n_code_links":2,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":0}},{"paper":"/paper/dilated-neighborhood-attention-transformer","title":"Dilated Neighborhood Attention Transformer","date":"2022-09-29","arxiv_id":"2209.15001","n_code_links":7,"syntology":null},{"paper":"/paper/v-2-l-leveraging-vision-and-vision-language","title":"V$^2$L: Leveraging Vision and Vision-language Models into Large-scale Product Retrieval","date":"2022-07-26","arxiv_id":"2207.12994","n_code_links":1,"syntology":null},{"paper":null,"title":"GAF-NAU: Gramian Angular Field encoded Neighborhood Attention U-Net for Pixel-Wise Hyperspectral Image Classification","date":"2022-04-21","arxiv_id":"2204.10099","n_code_links":0,"syntology":null},{"paper":"/paper/neighborhood-attention-transformer","title":"Neighborhood Attention Transformer","date":"2022-04-14","arxiv_id":"2204.07143","n_code_links":5,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}}],"papers_shown":21,"tasks":[{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":7},{"task":"/task/image-classification","name":"Image Classification","papers":4},{"task":"/task/decoder","name":"Decoder","papers":3},{"task":"/task/image-segmentation","name":"Image Segmentation","papers":3},{"task":"/task/object-detection","name":"Object Detection","papers":3},{"task":"/task/segmentation","name":"Segmentation","papers":3},{"task":"/task/image-classification","name":"image-classification","papers":3},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":2},{"task":"/task/denoising","name":"Denoising","papers":2},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":2},{"task":"/task/panoptic-segmentation","name":"Panoptic Segmentation","papers":2},{"task":"/task/ssim","name":"SSIM","papers":2},{"task":"/task/3d-object-detection","name":"3D Object Detection","papers":1},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":1},{"task":"/task/classification-1","name":"Classification","papers":1},{"task":"/task/deblurring","name":"Deblurring","papers":1},{"task":"/task/depth-estimation","name":"Depth Estimation","papers":1},{"task":"/task/diagnostic","name":"Diagnostic","papers":1},{"task":"/task/edge-detection","name":"Edge Detection","papers":1},{"task":"/task/emotion-recognition","name":"Emotion Recognition","papers":1}],"tasks_shown":20,"n_tasks":47,"usage_by_year":[{"year":"2022","papers":6},{"year":"2023","papers":5},{"year":"2024","papers":6},{"year":"2025","papers":4}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/na"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}