Papers › PVT v2: Improved Baselines with Pyramid Vision Transformer

PVT v2: Improved Baselines with Pyramid Vision Transformer

25 Jun 2021arXiv:2106.13797archive 2025-07-28

Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, Ling Shao

Transformer recently has presented encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs, including (1) linear complexity attention layer, (2) overlapping patch embedding, and (3) convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linear and achieves significant improvements on fundamental vision tasks such as classification, detection, and segmentation. Notably, the proposed PVT v2 achieves comparable or better performances than recent works such as Swin Transformer. We hope this work will facilitate state-of-the-art Transformer researches in computer vision. Code is available at https://github.com/whai362/PVT.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

18 repositories listed; official and paper-mentioned ones first.

whai362/PVT officialmentioned in papermentioned on GitHubpytorch report
Owais-Ansari/Unet3plus mentioned on GitHubpytorch report
open-mmlab/mmpose mentioned on GitHubpytorchApache-2.0 report
rwightman/pytorch-image-models mentioned on GitHubpytorch report
sithu31296/semantic-segmentation mentioned on GitHubpytorch report
xiaohu2015/pvt_detectron2 mentioned on GitHubpytorchMIT report
PaddlePaddle/PaddleClas paddleApache-2.0 report
open-mmlab/mmdetection pytorchApache-2.0 report
shinya7y/UniverseNet pytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image ClassificationObject DetectionPanoptic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ImageNet PVTv2-B4 GFLOPs 11.8 #391 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B4 Number of params 82M #391 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B4 Top 1 Accuracy 83.8% #391 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B3 GFLOPs 6.9 #457 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B3 Number of params 45.2M #457 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B3 Top 1 Accuracy 83.2% #457 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B2 GFLOPs 4 #585 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B2 Number of params 25.4M #585 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B2 Top 1 Accuracy 82% #585 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B1 GFLOPs 2.1 #813 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B1 Number of params 13.1M #813 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B1 Top 1 Accuracy 78.7% #813 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B0 GFLOPs 0.6 #1021 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B0 Number of params 3.4M #1021 of 1060 Archive leaderboard report
Image Classification ImageNet PVTv2-B0 Top 1 Accuracy 70.5% #1021 of 1060 Archive leaderboard report
Object Detection COCO minival Sparse R-CNN (PVTv2-B2) AP50 69.5 #80 of 220 Archive leaderboard report
Object Detection COCO minival Sparse R-CNN (PVTv2-B2) AP75 54.9 #80 of 220 Archive leaderboard report
Object Detection COCO minival Sparse R-CNN (PVTv2-B2) box AP 50.1 #80 of 220 Archive leaderboard report
Object Detection COCO-O PVTv2-B5 (Mask R-CNN) Average mAP 28.2 #23 of 45 Archive leaderboard report
Object Detection COCO-O PVTv2-B5 (Mask R-CNN) Effective Robustness 6.85 #23 of 45 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: PVTv2

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDepthwise ConvolutionDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPVTv2Position-Wise Feed-Forward LayerResidual ConnectionSoftmaxStochastic DepthSwin TransformerTransformerVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections