{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/agileformer-spatially-agile-transformer-unet","title":"AgileFormer: Spatially Agile Transformer UNet for Medical Image Segmentation","arxiv_id":"2404.00122","date":"2024-03-29","proceeding":null,"authors":["Peijie Qiu","Jin Yang","Sayantan Kumar","Soumyendu Sekhar Ghosh","Aristeidis Sotiras"],"abstract":"In the past decades, deep neural networks, particularly convolutional neural networks, have achieved state-of-the-art performance in a variety of medical image segmentation tasks. Recently, the introduction of the vision transformer (ViT) has significantly altered the landscape of deep segmentation models. There has been a growing focus on ViTs, driven by their excellent performance and scalability. However, we argue that the current design of the vision transformer-based UNet (ViT-UNet) segmentation models may not effectively handle the heterogeneous appearance (e.g., varying shapes and sizes) of objects of interest in medical image segmentation tasks. To tackle this challenge, we present a structured approach to introduce spatially dynamic components to the ViT-UNet. This adaptation enables the model to effectively capture features of target objects with diverse appearances. This is achieved by three main components: \\textbf{(i)} deformable patch embedding; \\textbf{(ii)} spatially dynamic multi-head attention; \\textbf{(iii)} deformable positional encoding. These components were integrated into a novel architecture, termed AgileFormer. AgileFormer is a spatially agile ViT-UNet designed for medical image segmentation. Experiments in three segmentation tasks using publicly available datasets demonstrated the effectiveness of the proposed method. The code is available at \\href{https://github.com/sotiraslab/AgileFormer}{https://github.com/sotiraslab/AgileFormer}.","url_abs":"https://arxiv.org/abs/2404.00122v2","url_pdf":"https://arxiv.org/pdf/2404.00122v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"agileformer-spatially-agile-transformer-unet","repo_url":"https://github.com/sotiraslab/AgileFormer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"medical-image-segmentation","task_name":"Medical Image Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unet-segmentation","task_name":"UNET Segmentation"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"focus","method_name":"Focus"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/medical-image-segmentation-on-acdc","task":"Medical Image Segmentation","dataset":"ACDC","model":"AgileFormer","rank_in_archive_order":2,"of":6,"metrics":{"Dice Score":"0.9255"},"uses_additional_data":false},{"leaderboard":"/sota/medical-image-segmentation-on-synapse-multi","task":"Medical Image Segmentation","dataset":"Synapse multi-organ CT","model":"AgileFormer","rank_in_archive_order":8,"of":23,"metrics":{"Avg DSC":"86.11","Avg HD":"12.88"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2404.00122","atlas_url":"https://app.syntology.ai/?focus=2404.00122","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}