{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transferring-to-real-world-layouts-a-depth","title":"Transferring to Real-World Layouts: A Depth-aware Framework for Scene Adaptation","arxiv_id":"2311.12682","date":"2023-11-21","proceeding":null,"authors":["Mu Chen","Zhedong Zheng","Yi Yang"],"abstract":"Scene segmentation via unsupervised domain adaptation (UDA) enables the transfer of knowledge acquired from source synthetic data to real-world target data, which largely reduces the need for manual pixel-level annotations in the target domain. To facilitate domain-invariant feature learning, existing methods typically mix data from both the source domain and target domain by simply copying and pasting the pixels. Such vanilla methods are usually sub-optimal since they do not take into account how well the mixed layouts correspond to real-world scenarios. Real-world scenarios are with an inherent layout. We observe that semantic categories, such as sidewalks, buildings, and sky, display relatively consistent depth distributions, and could be clearly distinguished in a depth map. Based on such observation, we propose a depth-aware framework to explicitly leverage depth estimation to mix the categories and facilitate the two complementary tasks, i.e., segmentation and depth learning in an end-to-end manner. In particular, the framework contains a Depth-guided Contextual Filter (DCF) forndata augmentation and a cross-task encoder for contextual learning. DCF simulates the real-world layouts, while the cross-task encoder further adaptively fuses the complementing features between two tasks. Besides, it is worth noting that several public datasets do not provide depth annotation. Therefore, we leverage the off-the-shelf depth estimation network to generate the pseudo depth. Extensive experiments show that our proposed methods, even with pseudo depth, achieve competitive performance on two widely-used bench-marks, i.e. 77.7 mIoU on GTA to Cityscapes and 69.3 mIoU on Synthia to Cityscapes.","url_abs":"https://arxiv.org/abs/2311.12682v1","url_pdf":"https://arxiv.org/pdf/2311.12682v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"transferring-to-real-world-layouts-a-depth","repo_url":"https://github.com/chen742/DCF","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"transferring-to-real-world-layouts-a-depth","repo_url":"https://github.com/chen742/PiPa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"domain-adaptation","task_name":"Domain Adaptation"},{"task_slug":"scene-segmentation","task_name":"Scene Segmentation"},{"task_slug":"synthetic-to-real-translation","task_name":"Synthetic-to-Real Translation"},{"task_slug":"unsupervised-domain-adaptation","task_name":"Unsupervised Domain Adaptation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/domain-adaptation-on-gta5-to-cityscapes","task":"Domain Adaptation","dataset":"GTA5 to Cityscapes","model":"DCF","rank_in_archive_order":2,"of":28,"metrics":{"mIoU":"77.7"},"uses_additional_data":false},{"leaderboard":"/sota/domain-adaptation-on-synthia-to-cityscapes","task":"Domain Adaptation","dataset":"SYNTHIA-to-Cityscapes","model":"DCF","rank_in_archive_order":3,"of":33,"metrics":{"mIoU":"69.3"},"uses_additional_data":false},{"leaderboard":"/sota/synthetic-to-real-translation-on-gtav-to","task":"Synthetic-to-Real Translation","dataset":"GTAV-to-Cityscapes Labels","model":"DCF","rank_in_archive_order":1,"of":73,"metrics":{"mIoU":"77.7"},"uses_additional_data":false},{"leaderboard":"/sota/synthetic-to-real-translation-on-synthia-to-1","task":"Synthetic-to-Real Translation","dataset":"SYNTHIA-to-Cityscapes","model":"DCF","rank_in_archive_order":1,"of":38,"metrics":{"MIoU (13 classes)":"75.9","MIoU (16 classes)":"69.3"},"uses_additional_data":true},{"leaderboard":"/sota/unsupervised-domain-adaptation-on-synthia-to","task":"Unsupervised Domain Adaptation","dataset":"SYNTHIA-to-Cityscapes","model":"DCF","rank_in_archive_order":1,"of":23,"metrics":{"MIoU (16 classes)":"69.3","mIoU":"69.3","mIoU (13 classes)":"75.9"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2311.12682","atlas_url":"https://app.syntology.ai/?focus=2311.12682","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}