Papers › Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers

Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers

16 Aug 2021arXiv:2108.06932archive 2025-07-28

Bo Dong, Wenhai Wang, Deng-Ping Fan, Jinpeng Li, Huazhu Fu, Ling Shao

Most polyp segmentation methods use CNNs as their backbone, leading to two key issues when exchanging information between the encoder and decoder: 1) taking into account the differences in contribution between different-level features and 2) designing an effective mechanism for fusing these features. Unlike existing CNN-based methods, we adopt a transformer encoder, which learns more powerful and robust representations. In addition, considering the image acquisition influence and elusive properties of polyps, we introduce three standard modules, including a cascaded fusion module (CFM), a camouflage identification module (CIM), and a similarity aggregation module (SAM). Among these, the CFM is used to collect the semantic and location information of polyps from high-level features; the CIM is applied to capture polyp information disguised in low-level features, and the SAM extends the pixel features of the polyp area with high-level semantic position information to the entire polyp area, thereby effectively fusing cross-level features. The proposed model, named Polyp-PVT, effectively suppresses noises in the features and significantly improves their expressive capabilities. Extensive experiments on five widely adopted datasets show that the proposed model is more robust to various challenging situations (e.g., appearance changes, small objects, rotation) than existing representative methods. The proposed model is available at https://github.com/DengPingFan/Polyp-PVT.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

DengPingFan/Polyp-PVT officialmentioned in papermentioned on GitHubpytorch report
whai362/PVT mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderMedical Image Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Medical Image Segmentation CVC-ColonDB Polyp-PVT Average MAE 0.031 #13 of 25 Archive leaderboard report
Medical Image Segmentation CVC-ColonDB Polyp-PVT S-Measure 0.865 #13 of 25 Archive leaderboard report
Medical Image Segmentation CVC-ColonDB Polyp-PVT mIoU 0.727 #13 of 25 Archive leaderboard report
Medical Image Segmentation CVC-ColonDB Polyp-PVT max E-Measure 0.913 #13 of 25 Archive leaderboard report
Medical Image Segmentation CVC-ColonDB Polyp-PVT mean Dice 0.808 #13 of 25 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections