{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mla-net-feature-pyramid-network-with-multi","title":"MLA-Net: Feature Pyramid Network with Multi-Level Local Attention for Object Detection","arxiv_id":null,"date":"2022-12-06","proceeding":"Mathematics 2022 12","authors":["Xiaobao Yang 1","2","*","Wentao Wang 3","Junsheng Wu 4","Chen Ding 3","Sugang Ma 3 and Zhiqiang Hou 3"],"abstract":"Abstract: Feature pyramid networks and attention mechanisms are the mainstream methods to\r\nimprove the detection performance of many current models. However, when they are learned\r\njointly, there is a lack of information association between multi-level features. Therefore, this paper\r\nproposes a feature pyramid of the multi-level local attention method, dubbed as MLA-Net (Feature\r\nPyramid Network with Multi-Level Local Attention for Object Detection), which aims to establish a\r\ncorrelation mechanism for multi-level local information. First, the original multi-level features are\r\ndeformed and rectified using the local pixel-rectification module, and global semantic enhancement is\r\nachieved through the multi-level spatial-attention module. After that, the original features are further\r\nfused through the residual connection to achieve the fusion of contextual features to enhance the\r\nfeature representation. Extensive ablation experiments were conducted on the MS COCO (Microsoft\r\nCommon Objects in Context) dataset, and the results demonstrate the effectiveness of the proposed\r\nmethod with a 0.5% enhancement. An improvement of 1.2% was obtained on the PASCAL VOC\r\n(Pattern Analysis Statistical Modelling and Computational Learning, Visual Object Classes) dataset,\r\nreaching 81.8%, thereby, indicating that the proposed method is robust and can compete with other\r\nadvanced detection models.","url_abs":"https://www.mdpi.com/2227-7390/10/24/4789/pdf?version=1671185236","url_pdf":"https://www.mdpi.com/2227-7390/10/24/4789/pdf?version=1671185236","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mla-net-feature-pyramid-network-with-multi","repo_url":"https://github.com/y78h11b09/mla-net","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}