Papers › AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation
AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation
Kashu Yamazaki, Taisei Hanyu, Minh Tran, Adrian de Luis, Roy McCann, Haitao Liao, Chase Rainwater, Meredith Adkins, Jackson Cothren, Ngan Le
Aerial Image Segmentation is a top-down perspective semantic segmentation and has several challenging characteristics such as strong imbalance in the foreground-background distribution, complex background, intra-class heterogeneity, inter-class homogeneity, and tiny objects. To handle these problems, we inherit the advantages of Transformers and propose AerialFormer, which unifies Transformers at the contracting path with lightweight Multi-Dilated Convolutional Neural Networks (MD-CNNs) at the expanding path. Our AerialFormer is designed as a hierarchical structure, in which Transformer encoder outputs multi-scale features and MD-CNNs decoder aggregates information from the multi-scales. Thus, it takes both local and global contexts into consideration to render powerful representations and high-resolution segmentation. We have benchmarked AerialFormer on three common datasets including iSAID, LoveDA, and Potsdam. Comprehensive experiments and extensive ablation studies show that our proposed AerialFormer outperforms previous state-of-the-art methods with remarkable performance. Our source code will be publicly available upon acceptance.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Semantic Segmentation | ISPRS Potsdam | AerialFormer-B | Mean F1 | 94.1 | #1 of 20 | Archive leaderboard | report |
| Semantic Segmentation | ISPRS Potsdam | AerialFormer-B | Mean IoU | 89.1 | #1 of 20 | Archive leaderboard | report |
| Semantic Segmentation | ISPRS Potsdam | AerialFormer-B | Overall Accuracy | 93.9 | #1 of 20 | Archive leaderboard | report |
| Semantic Segmentation | LoveDA | AerialFormer-B | Category mIoU | 54.1 | #8 of 19 | Archive leaderboard | report |
| Semantic Segmentation | iSAID | AerialFormer-B | mIoU | 69.3 | #3 of 19 | Archive leaderboard | report |
| Semantic Segmentation | iSAID | AerialFormer-S | mIoU | 68.4 | #5 of 19 | Archive leaderboard | report |
| Semantic Segmentation | iSAID | AerialFormer-T | mIoU | 67.5 | #9 of 19 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections