{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/zero-shot-visual-recognition-using-semantics","title":"Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks","arxiv_id":"1712.01928","date":"2017-12-05","proceeding":"CVPR 2018 6","authors":["Long Chen","Hanwang Zhang","Jun Xiao","Wei Liu","Shih-Fu Chang"],"abstract":"We propose a novel framework called Semantics-Preserving Adversarial\nEmbedding Network (SP-AEN) for zero-shot visual recognition (ZSL), where test\nimages and their classes are both unseen during training. SP-AEN aims to tackle\nthe inherent problem --- semantic loss --- in the prevailing family of\nembedding-based ZSL, where some semantics would be discarded during training if\nthey are non-discriminative for training classes, but could become critical for\nrecognizing test classes. Specifically, SP-AEN prevents the semantic loss by\nintroducing an independent visual-to-semantic space embedder which disentangles\nthe semantic space into two subspaces for the two arguably conflicting\nobjectives: classification and reconstruction. Through adversarial learning of\nthe two subspaces, SP-AEN can transfer the semantics from the reconstructive\nsubspace to the discriminative one, accomplishing the improved zero-shot\nrecognition of unseen classes. Comparing with prior works, SP-AEN can not only\nimprove classification but also generate photo-realistic images, demonstrating\nthe effectiveness of semantic preservation. On four popular benchmarks: CUB,\nAWA, SUN and aPY, SP-AEN considerably outperforms other state-of-the-art\nmethods by an absolute performance difference of 12.2\\%, 9.3\\%, 4.0\\%, and\n3.6\\% in terms of harmonic mean values","url_abs":"http://arxiv.org/abs/1712.01928v2","url_pdf":"http://arxiv.org/pdf/1712.01928v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"zero-shot-visual-recognition-using-semantics","repo_url":"https://github.com/MARMOTatZJU/ZSLPR-TIANCHI","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.01928","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}