{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-visual-representation-learning","title":"Unsupervised Visual Representation Learning by Context Prediction","arxiv_id":"1505.05192","date":"2015-05-19","proceeding":"ICCV 2015 12","authors":["Carl Doersch","Abhinav Gupta","Alexei A. Efros"],"abstract":"This work explores the use of spatial context as a source of free and\nplentiful supervisory signal for training a rich visual representation. Given\nonly a large, unlabeled image collection, we extract random pairs of patches\nfrom each image and train a convolutional neural net to predict the position of\nthe second patch relative to the first. We argue that doing well on this task\nrequires the model to learn to recognize objects and their parts. We\ndemonstrate that the feature representation learned using this within-image\ncontext indeed captures visual similarity across images. For example, this\nrepresentation allows us to perform unsupervised visual discovery of objects\nlike cats, people, and even birds from the Pascal VOC 2011 detection dataset.\nFurthermore, we show that the learned ConvNet can be used in the R-CNN\nframework and provides a significant boost over a randomly-initialized ConvNet,\nresulting in state-of-the-art performance among algorithms which use only\nPascal-provided training set annotations.","url_abs":"http://arxiv.org/abs/1505.05192v3","url_pdf":"http://arxiv.org/pdf/1505.05192v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-visual-representation-learning","repo_url":"https://github.com/open-mmlab/mmselfsup","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"unsupervised-visual-representation-learning","repo_url":"https://github.com/virtualgraham/sc_patch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"unsupervised-visual-representation-learning","repo_url":"https://github.com/abhisheksambyal/Self-supervised-learning-by-context-prediction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1505.05192","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}