Papers › Representation Separation for Semantic Segmentation with Vision Transformers

Representation Separation for Semantic Segmentation with Vision Transformers

28 Dec 2022arXiv:2212.13764archive 2025-07-28

Yuanduo Hong, Huihui Pan, Weichao Sun, Xinghu Yu, Huijun Gao

Vision transformers (ViTs) encoding an image as a sequence of patches bring new paradigms for semantic segmentation.We present an efficient framework of representation separation in local-patch level and global-region level for semantic segmentation with ViTs. It is targeted for the peculiar over-smoothness of ViTs in semantic segmentation, and therefore differs from current popular paradigms of context modeling and most existing related methods reinforcing the advantage of attention. We first deliver the decoupled two-pathway network in which another pathway enhances and passes down local-patch discrepancy complementary to global representations of transformers. We then propose the spatially adaptive separation module to obtain more separate deep representations and the discriminative cross-attention which yields more discriminative region representations through novel auxiliary supervisions. The proposed methods achieve some impressive results: 1) incorporated with large-scale plain ViTs, our methods achieve new state-of-the-art performances on five widely used benchmarks; 2) using masked pre-trained plain ViTs, we achieve 68.9% mIoU on Pascal Context, setting a new record; 3) pyramid ViTs integrated with the decoupled two-pathway network even surpass the well-designed high-resolution ViTs on Cityscapes; 4) the improved representations by our framework have favorable transferability in images with natural corruptions. The codes will be released publicly.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semantic Segmentation ADE20K RSSeg-ViT-L (BEiT pretrain) Params (M) 330 #20 of 235 Archive leaderboard report
Semantic Segmentation ADE20K RSSeg-ViT-L (BEiT pretrain) Validation mIoU 58.4 #20 of 235 Archive leaderboard report
Semantic Segmentation ADE20K val RSSeg-ViT-L(BEiT pretrain) mIoU 58.4 #12 of 95 Archive leaderboard report
Semantic Segmentation COCO-Stuff test RSSeg-ViT-L (BEiT pretrain) mIoU 52.6% #4 of 21 Archive leaderboard report
Semantic Segmentation COCO-Stuff test RSSeg-ViT-L mIoU 52.0% #5 of 21 Archive leaderboard report
Semantic Segmentation PASCAL Context RSSeg-ViT-L (BEiT pretrain) mIoU 68.9 #4 of 66 Archive leaderboard report
Semantic Segmentation PASCAL Context RSSeg-ViT-L mIoU 67.5 #7 of 66 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections