{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hierarchical-multi-scale-attention-for","title":"Hierarchical Multi-Scale Attention for Semantic Segmentation","arxiv_id":"2005.10821","date":"2020-05-21","proceeding":null,"authors":["Andrew Tao","Karan Sapra","Bryan Catanzaro"],"abstract":"Multi-scale inference is commonly used to improve the results of semantic segmentation. Multiple images scales are passed through a network and then the results are combined with averaging or max pooling. In this work, we present an attention-based approach to combining multi-scale predictions. We show that predictions at certain scales are better at resolving particular failures modes, and that the network learns to favor those scales for such cases in order to generate better predictions. Our attention mechanism is hierarchical, which enables it to be roughly 4x more memory efficient to train than other recent approaches. In addition to enabling faster training, this allows us to train with larger crop sizes which leads to greater model accuracy. We demonstrate the result of our method on two datasets: Cityscapes and Mapillary Vistas. For Cityscapes, which has a large number of weakly labelled images, we also leverage auto-labelling to improve generalization. Using our approach we achieve a new state-of-the-art results in both Mapillary (61.1 IOU val) and Cityscapes (85.1 IOU test).","url_abs":"https://arxiv.org/abs/2005.10821v1","url_pdf":"https://arxiv.org/pdf/2005.10821v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/NVIDIA/semantic-segmentation","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/Song-Jingyu/PointPainting","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/Song-Jingyu/runnable-PointPainting","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/YeLyuUT/semantic_segmentation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/ben-z-original/detectionhma","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/valeoai/SemanticPalette","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/JanMarcelKezmann/TensorFlow-Advanced-Segmentation-Models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null},{"paper_slug":"hierarchical-multi-scale-attention-for","repo_url":"https://github.com/PaddlePaddle/PaddleSeg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"paddle","reach":null}],"tasks":[{"task_slug":"panoptic-segmentation","task_name":"Panoptic Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/panoptic-segmentation-on-mapillary-val","task":"Panoptic Segmentation","dataset":"Mapillary val","model":"HRNet-OCR (Hierarchical Multi-Scale Attention)","rank_in_archive_order":12,"of":13,"metrics":{"PQ":"17.6"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-cityscapes-val","task":"Semantic Segmentation","dataset":"Cityscapes val","model":"HRNet-OCR","rank_in_archive_order":7,"of":99,"metrics":{"mIoU":"86.3"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2005.10821","atlas_url":"https://app.syntology.ai/?focus=2005.10821","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2005.10821"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/PaddlePaddle/PaddleSeg","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/NVIDIA/semantic-segmentation","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ben-z-original/detectionhma","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Song-Jingyu/runnable-PointPainting","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/JanMarcelKezmann/TensorFlow-Advanced-Segmentation-Models","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/valeoai/SemanticPalette","reach":{"status":"ok","spdx":"NOASSERTION"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Song-Jingyu/PointPainting","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/YeLyuUT/semantic_segmentation","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"summary":{"ran_draft_wrong":1,"ran":1,"unverified":7},"by_repo_kind":{"official":{"samples":5,"ran":1,"repositories":1},"listed":{"samples":4,"ran":1,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"fac5364e2f53c6db","entry":"conv3x3","repo":"NVIDIA/semantic-segmentation","repo_kind":"official","path":"network/Resnet.py","file_url":"https://github.com/NVIDIA/semantic-segmentation/blob/HEAD/network/Resnet.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"fac5364e2f53c6db"}},{"code_sha256_prefix":"fe4acbd5f88b4231","entry":"get_valid_ratio","repo":"YeLyuUT/semantic_segmentation","repo_kind":"listed","path":"network/adnb.py","file_url":"https://github.com/YeLyuUT/semantic_segmentation/blob/HEAD/network/adnb.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":false,"mcp_get_code":{"code_sha256":"fe4acbd5f88b4231"}},{"code_sha256_prefix":"696fe155f9990d38","entry":"cfg_from_yaml_file","repo":"Song-Jingyu/PointPainting","repo_kind":"listed","path":"detector/pcdet/config.py","file_url":"https://github.com/Song-Jingyu/PointPainting/blob/HEAD/detector/pcdet/config.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"696fe155f9990d38"}},{"code_sha256_prefix":"853a00663733a69a","entry":"customsoftmax","repo":"NVIDIA/semantic-segmentation","repo_kind":"official","path":"loss/utils.py","file_url":"https://github.com/NVIDIA/semantic-segmentation/blob/HEAD/loss/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"853a00663733a69a"}},{"code_sha256_prefix":"1780d388cc532a6d","entry":"get_corner_loss_lidar","repo":"Song-Jingyu/PointPainting","repo_kind":"listed","path":"detector/pcdet/utils/loss_utils.py","file_url":"https://github.com/Song-Jingyu/PointPainting/blob/HEAD/detector/pcdet/utils/loss_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1780d388cc532a6d"}},{"code_sha256_prefix":"c48d876b9a3ae27a","entry":"log_det_by_cholesky","repo":"NVIDIA/semantic-segmentation","repo_kind":"official","path":"loss/rmi_utils.py","file_url":"https://github.com/NVIDIA/semantic-segmentation/blob/HEAD/loss/rmi_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"c48d876b9a3ae27a"}},{"code_sha256_prefix":"97f5e3b69f2c69ea","entry":"map_get_pairs","repo":"NVIDIA/semantic-segmentation","repo_kind":"official","path":"loss/rmi_utils.py","file_url":"https://github.com/NVIDIA/semantic-segmentation/blob/HEAD/loss/rmi_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"97f5e3b69f2c69ea"}},{"code_sha256_prefix":"99b42d109e83351a","entry":"map_get_pairs_region","repo":"NVIDIA/semantic-segmentation","repo_kind":"official","path":"loss/rmi_utils.py","file_url":"https://github.com/NVIDIA/semantic-segmentation/blob/HEAD/loss/rmi_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"99b42d109e83351a"}},{"code_sha256_prefix":"392c0cf3a1b07b12","entry":"merge_new_config","repo":"Song-Jingyu/PointPainting","repo_kind":"listed","path":"detector/pcdet/config.py","file_url":"https://github.com/Song-Jingyu/PointPainting/blob/HEAD/detector/pcdet/config.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"392c0cf3a1b07b12"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}