{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/texttopicnet-self-supervised-learning-of","title":"TextTopicNet - Self-Supervised Learning of Visual Features Through Embedding Images on Semantic Text Spaces","arxiv_id":"1807.02110","date":"2018-07-04","proceeding":null,"authors":["Yash Patel","Lluis Gomez","Raul Gomez","Marçal Rusiñol","Dimosthenis Karatzas","C. V. Jawahar"],"abstract":"The immense success of deep learning based methods in computer vision heavily\nrelies on large scale training datasets. These richly annotated datasets help\nthe network learn discriminative visual features. Collecting and annotating\nsuch datasets requires a tremendous amount of human effort and annotations are\nlimited to popular set of classes. As an alternative, learning visual features\nby designing auxiliary tasks which make use of freely available\nself-supervision has become increasingly popular in the computer vision\ncommunity.\n  In this paper, we put forward an idea to take advantage of multi-modal\ncontext to provide self-supervision for the training of computer vision\nalgorithms. We show that adequate visual features can be learned efficiently by\ntraining a CNN to predict the semantic textual context in which a particular\nimage is more probable to appear as an illustration. More specifically we use\npopular text embedding techniques to provide the self-supervision for the\ntraining of deep CNN.\n  Our experiments demonstrate state-of-the-art performance in image\nclassification, object detection, and multi-modal retrieval compared to recent\nself-supervised or naturally-supervised approaches.","url_abs":"http://arxiv.org/abs/1807.02110v1","url_pdf":"http://arxiv.org/pdf/1807.02110v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"texttopicnet-self-supervised-learning-of","repo_url":"https://github.com/lluisgomez/TextTopicNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"image-classification","task_name":"image-classification"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}