{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-robust-visual-semantic-embeddings","title":"Learning Robust Visual-Semantic Embeddings","arxiv_id":"1703.05908","date":"2017-03-17","proceeding":"ICCV 2017 10","authors":["Yao-Hung Hubert Tsai","Liang-Kang Huang","Ruslan Salakhutdinov"],"abstract":"Many of the existing methods for learning joint embedding of images and text\nuse only supervised information from paired images and its textual attributes.\nTaking advantage of the recent success of unsupervised learning in deep neural\nnetworks, we propose an end-to-end learning framework that is able to extract\nmore robust multi-modal representations across domains. The proposed method\ncombines representation learning models (i.e., auto-encoders) together with\ncross-domain learning criteria (i.e., Maximum Mean Discrepancy loss) to learn\njoint embeddings for semantic and visual features. A novel technique of\nunsupervised-data adaptation inference is introduced to construct more\ncomprehensive embeddings for both labeled and unlabeled data. We evaluate our\nmethod on Animals with Attributes and Caltech-UCSD Birds 200-2011 dataset with\na wide range of applications, including zero and few-shot image recognition and\nretrieval, from inductive to transductive settings. Empirically, we show that\nour framework improves over the current state of the art on many of the\nconsidered tasks.","url_abs":"http://arxiv.org/abs/1703.05908v2","url_pdf":"http://arxiv.org/pdf/1703.05908v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"generalized-few-shot-learning","task_name":"Generalized Few-Shot Learning"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/generalized-few-shot-learning-on-awa2","task":"Generalized Few-Shot Learning","dataset":"AwA2","model":"REVISE","rank_in_archive_order":6,"of":6,"metrics":{"Per-Class Accuracy (1-shot)":"56.1","Per-Class Accuracy (10-shots)":"67.8","Per-Class Accuracy (2-shots)":"60.3","Per-Class Accuracy (5-shots)":"64.1"},"uses_additional_data":false},{"leaderboard":"/sota/generalized-few-shot-learning-on-cub","task":"Generalized Few-Shot Learning","dataset":"CUB","model":"REVISE","rank_in_archive_order":5,"of":5,"metrics":{"Per-Class Accuracy  (2-shots)":"41.1","Per-Class Accuracy (1-shot)":"36.3","Per-Class Accuracy (10-shots)":"50.9","Per-Class Accuracy (5-shots)":"44.6"},"uses_additional_data":false},{"leaderboard":"/sota/generalized-few-shot-learning-on-sun","task":"Generalized Few-Shot Learning","dataset":"SUN","model":"REVISE","rank_in_archive_order":5,"of":5,"metrics":{"Per-Class Accuracy (1-shot)":"27.4","Per-Class Accuracy (10-shots)":"40.8","Per-Class Accuracy (2-shots)":"33.4","Per-Class Accuracy (5-shots)":"37.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.05908","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}