Methods › Computer Vision › Vision Transformers › PVT
Pyramid Vision Transformer
PVT
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
PVT, or Pyramid Vision Transformer, is a type of vision transformer that utilizes a pyramid structure to make it an effective backbone for dense prediction tasks. Specifically it allows for more fine-grained inputs (4 x 4 pixels per patch) to be used, while simultaneously shrinking the sequence length of the Transformer as it deepens - reducing the computational cost. Additionally, a spatial-reduction attention (SRA) layer is used to further reduce the resource consumption when learning high-resolution features.
The entire model is divided into four stages, each of which is comprised of a patch embedding layer and a ℒᵢ-layer Transformer encoder. Following a pyramid structure, the output resolution of the four stages progressively shrinks from high (4-stride) to low (32-stride).
Papers archive 2025-07-28
28 shown of 28, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
GLOVA: Global and Local Variation-Aware Analog Circuit Design with Risk-Sensitive Reinforcement Learning 16 May 2025 · 0 repositories · arXiv:2505.11208
-
Crystal Oscillators in OSNMA-Enabled Receivers: An Implementation View for Automotive Applications 25 Jan 2025 · 0 repositories · arXiv:2501.15123
-
Multipath Mitigation Technology-integrated GNSS Direct Position Estimation Plug-in Module 20 Nov 2024 · 0 repositories · arXiv:2411.13339
-
HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation 29 Oct 2024 · 0 repositories · arXiv:2410.22079
-
Low-Rank Continual Pyramid Vision Transformer: Incrementally Segment Whole-Body Organs in CT with Light-Weighted Adaptation 7 Oct 2024 · 0 repositories · arXiv:2410.04689
-
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies 24 May 2024 · 0 repositories · arXiv:2405.15916
-
Rethinking Attention Gated with Hybrid Dual Pyramid Transformer-CNN for Generalized Segmentation in Medical Imaging 28 Apr 2024 · 1 repository · arXiv:2404.18199
-
Multi-Layer Dense Attention Decoder for Polyp Segmentation 27 Mar 2024 · 1 repository · arXiv:2403.18180
-
Heracles: A Hybrid SSM-Transformer Model for High-Resolution Image and Time-Series Analysis 26 Mar 2024 · 2 repositories · arXiv:2403.18063Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)
-
ROI-Aware Multiscale Cross-Attention Vision Transformer for Pest Image Identification 28 Dec 2023 · 0 repositories · arXiv:2312.16914
-
Distilling Knowledge from CNN-Transformer Models for Enhanced Human Action Recognition 2 Nov 2023 · 0 repositories · arXiv:2311.01283
-
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention 4 Sep 2023 · 1 repository · arXiv:2309.01430Syntology ran 2 of 10 samples · 8 unverified
-
A denoised Mean Teacher for domain adaptive point cloud registration 26 Jun 2023 · 1 repository · arXiv:2306.14749
-
A 3-step Low-latency Low-Power Multichannel Time-to-Digital Converter based on Time Residual Amplifier 1 Jun 2023 · 0 repositories · arXiv:2306.00433
-
Neural correlates of cognitive ability and visuo-motor speed: validation of IDoCT on UK Biobank Data 30 May 2023 · 0 repositories · arXiv:2305.18804
-
PVT-SSD: Single-Stage 3D Object Detector with Point-Voxel Transformer 11 May 2023 · 1 repository · arXiv:2305.06621
-
Sector Patch Embedding: An Embedding Module Conforming to The Distortion Pattern of Fisheye Image 26 Mar 2023 · 0 repositories · arXiv:2303.14645
-
Chasing Clouds: Differentiable Volumetric Rasterisation of Point Clouds as a Highly Efficient and Accurate Loss for Large-Scale Deformable 3D Registration 1 Jan 2023 · 1 repository
-
Exploring the Relationship Between Architectural Design and Adversarially Robust Generalization 1 Jan 2023 · 0 repositories
-
EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual Backbones 17 Nov 2022 · 1 repository · arXiv:2211.09703
-
Exploring the Relationship between Architecture and Adversarially Robust Generalization 28 Sep 2022 · 0 repositories · arXiv:2209.14105
-
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality 20 May 2022 · 1 repository · arXiv:2205.10063
-
Uncertainty-Cognizant Model Predictive Control for Energy Management of Residential Buildings with PVT and Thermal Energy Storage 21 Jan 2022 · 0 repositories · arXiv:2201.08909
-
Vision Transformer with Deformable Attention 3 Jan 2022 · 2 repositories · arXiv:2201.00520Syntology ran 8 of 10 samples · 2 unverified
-
Dynamic Token Normalization Improves Vision Transformers 5 Dec 2021 · 1 repository · arXiv:2112.02624Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
Partial Variable Training for Efficient On-Device Federated Learning 11 Oct 2021 · 0 repositories · arXiv:2110.05607
-
PVT: Point-Voxel Transformer for Point Cloud Learning 13 Aug 2021 · 2 repositories · arXiv:2108.06076
-
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions 24 Feb 2021 · 11 repositories · arXiv:2102.12122Syntology ran 22 of 30 samples · 8 unverified · 1 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 42 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections