{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pad-net-multi-tasks-guided-prediction-and","title":"PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing","arxiv_id":"1805.04409","date":"2018-05-11","proceeding":"CVPR 2018 6","authors":["Dan Xu","Wanli Ouyang","Xiaogang Wang","Nicu Sebe"],"abstract":"Depth estimation and scene parsing are two particularly important tasks in\nvisual scene understanding. In this paper we tackle the problem of simultaneous\ndepth estimation and scene parsing in a joint CNN. The task can be typically\ntreated as a deep multi-task learning problem [42]. Different from previous\nmethods directly optimizing multiple tasks given the input training data, this\npaper proposes a novel multi-task guided prediction-and-distillation network\n(PAD-Net), which first predicts a set of intermediate auxiliary tasks ranging\nfrom low level to high level, and then the predictions from these intermediate\nauxiliary tasks are utilized as multi-modal input via our proposed multi-modal\ndistillation modules for the final tasks. During the joint learning, the\nintermediate tasks not only act as supervision for learning more robust deep\nrepresentations but also provide rich multi-modal information for improving the\nfinal tasks. Extensive experiments are conducted on two challenging datasets\n(i.e. NYUD-v2 and Cityscapes) for both the depth estimation and scene parsing\ntasks, demonstrating the effectiveness of the proposed approach.","url_abs":"http://arxiv.org/abs/1805.04409v1","url_pdf":"http://arxiv.org/pdf/1805.04409v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"scene-parsing","task_name":"Scene Parsing"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/depth-estimation-on-nyu-depth-v2","task":"Depth Estimation","dataset":"NYU-Depth V2","model":"PAD-Net","rank_in_archive_order":15,"of":17,"metrics":{"RMS":"0.792"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.04409","atlas_url":"https://app.syntology.ai/?focus=1805.04409","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}