{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-modal-subspace-learning-for-fine","title":"Cross-modal Subspace Learning for Fine-grained Sketch-based Image Retrieval","arxiv_id":"1705.09888","date":"2017-05-28","proceeding":null,"authors":["Peng Xu","Qiyue Yin","Yongye Huang","Yi-Zhe Song","Zhanyu Ma","Liang Wang","Tao Xiang","W. Bastiaan Kleijn","Jun Guo"],"abstract":"Sketch-based image retrieval (SBIR) is challenging due to the inherent\ndomain-gap between sketch and photo. Compared with pixel-perfect depictions of\nphotos, sketches are iconic renderings of the real world with highly abstract.\nTherefore, matching sketch and photo directly using low-level visual clues are\nunsufficient, since a common low-level subspace that traverses semantically\nacross the two modalities is non-trivial to establish. Most existing SBIR\nstudies do not directly tackle this cross-modal problem. This naturally\nmotivates us to explore the effectiveness of cross-modal retrieval methods in\nSBIR, which have been applied in the image-text matching successfully. In this\npaper, we introduce and compare a series of state-of-the-art cross-modal\nsubspace learning methods and benchmark them on two recently released\nfine-grained SBIR datasets. Through thorough examination of the experimental\nresults, we have demonstrated that the subspace learning can effectively model\nthe sketch-photo domain-gap. In addition we draw a few key insights to drive\nfuture research.","url_abs":"http://arxiv.org/abs/1705.09888v1","url_pdf":"http://arxiv.org/pdf/1705.09888v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"image-text-matching","task_name":"Image-text matching"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"sketch-based-image-retrieval","task_name":"Sketch-Based Image Retrieval"},{"task_slug":"text-matching","task_name":"Text Matching"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/sketch-based-image-retrieval-on-chairs","task":"Sketch-Based Image Retrieval","dataset":"Chairs","model":"CCA-3V-HOG + PCA","rank_in_archive_order":5,"of":8,"metrics":{"R@1":"53.2","R@10":"90.3"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}