{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-task-self-supervised-visual-learning","title":"Multi-task Self-Supervised Visual Learning","arxiv_id":"1708.07860","date":"2017-08-25","proceeding":"ICCV 2017 10","authors":["Carl Doersch","Andrew Zisserman"],"abstract":"We investigate methods for combining multiple self-supervised tasks--i.e.,\nsupervised tasks where data can be collected without manual labeling--in order\nto train a single visual representation. First, we provide an apples-to-apples\ncomparison of four different self-supervised tasks using the very deep\nResNet-101 architecture. We then combine tasks to jointly train a network. We\nalso explore lasso regularization to encourage the network to factorize the\ninformation in its representation, and methods for \"harmonizing\" network inputs\nin order to learn a more unified representation. We evaluate all methods on\nImageNet classification, PASCAL VOC detection, and NYU depth prediction. Our\nresults show that deeper networks work better, and that combining tasks--even\nvia a naive multi-head architecture--always improves performance. Our best\njoint network nearly matches the PASCAL performance of a model pre-trained on\nImageNet classification, and matches the ImageNet network on NYU depth\nprediction.","url_abs":"http://arxiv.org/abs/1708.07860v1","url_pdf":"http://arxiv.org/pdf/1708.07860v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"depth-prediction","task_name":"Depth Prediction"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"self-supervised-image-classification","task_name":"Self-Supervised Image Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/self-supervised-image-classification-on","task":"Self-Supervised Image Classification","dataset":"ImageNet","model":"Colorisation (improved) (ResNet-101)","rank_in_archive_order":139,"of":144,"metrics":{"Number of Params":"44M","Top 1 Accuracy":"39.6","Top 5 Accuracy":"62.5"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-image-classification-on","task":"Self-Supervised Image Classification","dataset":"ImageNet","model":"Multi-task SSL (ResNet-101)","rank_in_archive_order":144,"of":144,"metrics":{"Number of Params":"44M","Top 5 Accuracy":"70.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.07860","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}