{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mtanet-multitask-aware-network-with","title":"MTANet: Multitask-Aware Network With Hierarchical Multimodal Fusion for RGB-T Urban Scene Understanding","arxiv_id":null,"date":"2022-04-05","proceeding":"journal 2022 4","authors":["WuJie Zhou","Shaohua Dong","Jingsheng Lei","Lu Yu"],"abstract":"Understanding urban scenes is a fundamental ability\r\nrequirement for assisted driving and autonomous vehicles. Most of\r\nthe available urban scene understanding methods use red-greenblue (RGB) images; however, their segmentation performances are\r\nprone to degradation under adverse lighting conditions. Recently,\r\nmany effective artificial neural networks have been presented for\r\nurban scene understanding and have shown that incorporating\r\nRGB and thermal (RGB-T) images can improve segmentation accuracy even under unsatisfactory lighting conditions. However, the\r\npotential of multimodal feature fusion has not been fully exploited\r\nbecause operations such as simply concatenating the RGB and\r\nthermal features or averaging their maps have been adopted. To\r\nimprove the fusion of multimodal features and the segmentation\r\naccuracy, we propose a multitask-aware network (MTANet) with\r\nhierarchical multimodal fusion (multiscale fusion strategy) for\r\nRGB-T urban scene understanding. We developed a hierarchical\r\nmultimodal fusion module to enhance feature fusion and built a\r\nhigh-level semantic module to extract semantic information for\r\nmerging with coarse features at various abstraction levels. Using the\r\nmultilevel fusion module, we exploited low-, mid-, and high-level\r\nfusion to improve segmentation accuracy. The multitask module\r\nuses boundary, binary, and semantic supervision to optimize the\r\nMTANet parameters. Extensive experiments were performed on\r\ntwo benchmark RGB-T datasets to verify the improved performance of the proposed MTANet compared with state-of-the-art\r\nmethods","url_abs":"https://ieeexplore.ieee.org/abstract/document/9749834","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9749834","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"autonomous-vehicles","task_name":"Autonomous Vehicles"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"thermal-image-segmentation","task_name":"Thermal Image Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/thermal-image-segmentation-on-mfn-dataset","task":"Thermal Image Segmentation","dataset":"MFN Dataset","model":"MTANet","rank_in_archive_order":27,"of":55,"metrics":{"mIOU":"56.1"},"uses_additional_data":false},{"leaderboard":"/sota/thermal-image-segmentation-on-pst900","task":"Thermal Image Segmentation","dataset":"PST900","model":"MTANet","rank_in_archive_order":14,"of":22,"metrics":{"mIoU":"78.60"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}