{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-media-similarity-evaluation-for-web","title":"Cross-Media Similarity Evaluation for Web Image Retrieval in the Wild","arxiv_id":"1709.01305","date":"2017-09-05","proceeding":null,"authors":["Jianfeng Dong","Xirong Li","Duanqing Xu"],"abstract":"In order to retrieve unlabeled images by textual queries, cross-media\nsimilarity computation is a key ingredient. Although novel methods are\ncontinuously introduced, little has been done to evaluate these methods\ntogether with large-scale query log analysis. Consequently, how far have these\nmethods brought us in answering real-user queries is unclear. Given baseline\nmethods that compute cross-media similarity using relatively simple text/image\nmatching, how much progress have advanced models made is also unclear. This\npaper takes a pragmatic approach to answering the two questions. Queries are\nautomatically categorized according to the proposed query visualness measure,\nand later connected to the evaluation of multiple cross-media similarity models\non three test sets. Such a connection reveals that the success of the\nstate-of-the-art is mainly attributed to their good performance on\nvisual-oriented queries, while these queries account for only a small part of\nreal-user queries. To quantify the current progress, we propose a simple\ntext2image method, representing a novel test query by a set of images selected\nfrom large-scale query log. Consequently, computing cross-media similarity\nbetween the test query and a given image boils down to comparing the visual\nsimilarity between the given image and the selected images. Image retrieval\nexperiments on the challenging Clickture dataset show that the proposed\ntext2image compares favorably to recent deep learning based alternatives.","url_abs":"http://arxiv.org/abs/1709.01305v2","url_pdf":"http://arxiv.org/pdf/1709.01305v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cross-media-similarity-evaluation-for-web","repo_url":"https://github.com/danieljf24/text2image","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.01305","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}