{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/u-net-ensemble-for-enhanced-semantic","title":"U-Net Ensemble for Enhanced Semantic Segmentation in Remote Sensing Imagery","arxiv_id":null,"date":"2024-06-08","proceeding":"Remote Sensing 2024 6","authors":["Ivica Dimitrovski","Vlatko Spasev","Suzana Loshkovska","Ivan Kitanovski"],"abstract":"Semantic segmentation of remote sensing imagery stands as a fundamental task within the domains of both remote sensing and computer vision. Its objective is to generate a comprehensive pixel-wise segmentation map of an image, assigning a specific label to each pixel. This facilitates in-depth analysis and comprehension of the Earth’s surface. In this paper, we propose an approach for enhancing semantic segmentation performance by employing an ensemble of U-Net models with three different backbone networks: Multi-Axis Vision Transformer, ConvFormer, and EfficientNet. The final segmentation maps are generated through a geometric mean ensemble method, leveraging the diverse representations learned by each backbone network. The effectiveness of the base U-Net models and the proposed ensemble is evaluated on multiple datasets commonly used for semantic segmentation tasks in remote sensing imagery, including LandCover.ai, LoveDA, INRIA, UAVid, and ISPRS Potsdam datasets. Our experimental results demonstrate that the proposed approach achieves state-of-the-art performance, showcasing its effectiveness and robustness in accurately capturing the semantic information embedded within remote sensing images.","url_abs":"https://www.mdpi.com/2072-4292/16/12/2077","url_pdf":"https://www.mdpi.com/2072-4292/16/12/2077/pdf?version=1717819425","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"segmentation-of-remote-sensing-imagery","task_name":"Segmentation Of Remote Sensing Imagery"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"base","method_name":"BASE"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"concatenated-skip-connection","method_name":"Concatenated Skip Connection"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"depthwise-convolution","method_name":"Depthwise Convolution"},{"method_slug":"depthwise-separable-convolution","method_name":"Depthwise Separable Convolution"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"efficientnet","method_name":"EfficientNet"},{"method_slug":"inverted-residual-block","method_name":"Inverted Residual Block"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"pointwise-convolution","method_name":"Pointwise Convolution"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"rmsprop","method_name":"RMSProp"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"squeeze-and-excitation-block","method_name":"Squeeze-and-Excitation Block"},{"method_slug":"transformer","method_name":"Transformer"},{"method_slug":"u-net","method_name":"U-Net"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-isprs-potsdam","task":"Semantic Segmentation","dataset":"ISPRS Potsdam","model":"U-Net (ConvFormer-M36)","rank_in_archive_order":20,"of":20,"metrics":{"Mean IoU":"89.45"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-landcover-ai","task":"Semantic Segmentation","dataset":"LandCover.ai","model":"U-Net (ConvFormer-M36)","rank_in_archive_order":1,"of":1,"metrics":{"mIoU":"87.64"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-loveda","task":"Semantic Segmentation","dataset":"LoveDA","model":"U-Net (MaxViT-S)","rank_in_archive_order":1,"of":19,"metrics":{"Category mIoU":"56.16"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-uavid","task":"Semantic Segmentation","dataset":"UAVid","model":"U-Net Ensemble","rank_in_archive_order":1,"of":10,"metrics":{"Mean IoU":"73.34"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-uavid","task":"Semantic Segmentation","dataset":"UAVid","model":"U-Net (MaxViT-S)","rank_in_archive_order":2,"of":10,"metrics":{"Mean IoU":"71.88"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}