{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/automatic-spatially-aware-fashion-concept","title":"Automatic Spatially-aware Fashion Concept Discovery","arxiv_id":"1708.01311","date":"2017-08-03","proceeding":"ICCV 2017 10","authors":["Xintong Han","Zuxuan Wu","Phoenix X. Huang","Xiao Zhang","Menglong Zhu","Yuan Li","Yang Zhao","Larry S. Davis"],"abstract":"This paper proposes an automatic spatially-aware concept discovery approach\nusing weakly labeled image-text data from shopping websites. We first fine-tune\nGoogleNet by jointly modeling clothing images and their corresponding\ndescriptions in a visual-semantic embedding space. Then, for each attribute\n(word), we generate its spatially-aware representation by combining its\nsemantic word vector representation with its spatial representation derived\nfrom the convolutional maps of the fine-tuned network. The resulting\nspatially-aware representations are further used to cluster attributes into\nmultiple groups to form spatially-aware concepts (e.g., the neckline concept\nmight consist of attributes like v-neck, round-neck, etc). Finally, we\ndecompose the visual-semantic embedding space into multiple concept-specific\nsubspaces, which facilitates structured browsing and attribute-feedback product\nretrieval by exploiting multimodal linguistic regularities. We conducted\nextensive experiments on our newly collected Fashion200K dataset, and results\non clustering quality evaluation and attribute-feedback product retrieval task\ndemonstrate the effectiveness of our automatically discovered spatially-aware\nconcepts.","url_abs":"http://arxiv.org/abs/1708.01311v1","url_pdf":"http://arxiv.org/pdf/1708.01311v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"automatic-spatially-aware-fashion-concept","repo_url":"https://github.com/naver/artemis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"multi-modal","task_name":"Image Retrieval with Multi-Modal Query"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-retrieval-with-multi-modal-query-on","task":"Image Retrieval with Multi-Modal Query","dataset":"Fashion200k","model":"FashionConcept","rank_in_archive_order":8,"of":8,"metrics":{"Recall@1":"6.3","Recall@10":"19.9","Recall@50":"38.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.01311","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}