{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ubernet-training-a-universal-convolutional","title":"UberNet: Training a `Universal' Convolutional Neural Network for Low-, Mid-, and High-Level Vision using Diverse Datasets and Limited Memory","arxiv_id":"1609.02132","date":"2016-09-07","proceeding":null,"authors":["Iasonas Kokkinos"],"abstract":"In this work we introduce a convolutional neural network (CNN) that jointly\nhandles low-, mid-, and high-level vision tasks in a unified architecture that\nis trained end-to-end. Such a universal network can act like a `swiss knife'\nfor vision tasks; we call this architecture an UberNet to indicate its\noverarching nature.\n  We address two main technical challenges that emerge when broadening up the\nrange of tasks handled by a single CNN: (i) training a deep architecture while\nrelying on diverse training sets and (ii) training many (potentially unlimited)\ntasks with a limited memory budget. Properly addressing these two problems\nallows us to train accurate predictors for a host of tasks, without\ncompromising accuracy.\n  Through these advances we train in an end-to-end manner a CNN that\nsimultaneously addresses (a) boundary detection (b) normal estimation (c)\nsaliency estimation (d) semantic segmentation (e) human part segmentation (f)\nsemantic boundary detection, (g) region proposal generation and object\ndetection. We obtain competitive performance while jointly addressing all of\nthese tasks in 0.7 seconds per frame on a single GPU. A demonstration of this\nsystem can be found at http://cvn.ecp.fr/ubernet/.","url_abs":"http://arxiv.org/abs/1609.02132v1","url_pdf":"http://arxiv.org/pdf/1609.02132v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ubernet-training-a-universal-convolutional","repo_url":"https://github.com/EPFL-VILAB/XDEnsembles","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"boundary-detection","task_name":"Boundary Detection"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"human-part-segmentation","task_name":"Human Part Segmentation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"region-proposal","task_name":"Region Proposal"},{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1609.02132","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}