Papers › AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation

AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation

12 Jun 2023arXiv:2306.06842archive 2025-07-28

Kashu Yamazaki, Taisei Hanyu, Minh Tran, Adrian de Luis, Roy McCann, Haitao Liao, Chase Rainwater, Meredith Adkins, Jackson Cothren, Ngan Le

Aerial Image Segmentation is a top-down perspective semantic segmentation and has several challenging characteristics such as strong imbalance in the foreground-background distribution, complex background, intra-class heterogeneity, inter-class homogeneity, and tiny objects. To handle these problems, we inherit the advantages of Transformers and propose AerialFormer, which unifies Transformers at the contracting path with lightweight Multi-Dilated Convolutional Neural Networks (MD-CNNs) at the expanding path. Our AerialFormer is designed as a hierarchical structure, in which Transformer encoder outputs multi-scale features and MD-CNNs decoder aggregates information from the multi-scales. Thus, it takes both local and global contexts into consideration to render powerful representations and high-resolution segmentation. We have benchmarked AerialFormer on three common datasets including iSAID, LoveDA, and Potsdam. Comprehensive experiments and extensive ablation studies show that our proposed AerialFormer outperforms previous state-of-the-art methods with remarkable performance. Our source code will be publicly available upon acceptance.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

UARK-AICV/AerialFormer officialmentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderImage SegmentationSegmentationSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semantic Segmentation ISPRS Potsdam AerialFormer-B Mean F1 94.1 #1 of 20 Archive leaderboard report
Semantic Segmentation ISPRS Potsdam AerialFormer-B Mean IoU 89.1 #1 of 20 Archive leaderboard report
Semantic Segmentation ISPRS Potsdam AerialFormer-B Overall Accuracy 93.9 #1 of 20 Archive leaderboard report
Semantic Segmentation LoveDA AerialFormer-B Category mIoU 54.1 #8 of 19 Archive leaderboard report
Semantic Segmentation iSAID AerialFormer-B mIoU 69.3 #3 of 19 Archive leaderboard report
Semantic Segmentation iSAID AerialFormer-S mIoU 68.4 #5 of 19 Archive leaderboard report
Semantic Segmentation iSAID AerialFormer-T mIoU 67.5 #9 of 19 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections