{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/composing-text-and-image-for-image-retrieval","title":"Composing Text and Image for Image Retrieval - An Empirical Odyssey","arxiv_id":"1812.07119","date":"2018-12-18","proceeding":"CVPR 2019 6","authors":["Nam Vo","Lu Jiang","Chen Sun","Kevin Murphy","Li-Jia Li","Li Fei-Fei","James Hays"],"abstract":"In this paper, we study the task of image retrieval, where the input query is\nspecified in the form of an image plus some text that describes desired\nmodifications to the input image. For example, we may present an image of the\nEiffel tower, and ask the system to find images which are visually similar but\nare modified in small ways, such as being taken at nighttime instead of during\nthe day. To tackle this task, we learn a similarity metric between a target\nimage and a source image plus source text, an embedding and composing function\nsuch that target image feature is close to the source image plus text\ncomposition feature. We propose a new way to combine image and text using such\nfunction that is designed for the retrieval task. We show this outperforms\nexisting approaches on 3 different datasets, namely Fashion-200k, MIT-States\nand a new synthetic dataset we create based on CLEVR. We also show that our\napproach can be used to classify input queries, in addition to image retrieval.","url_abs":"http://arxiv.org/abs/1812.07119v1","url_pdf":"http://arxiv.org/pdf/1812.07119v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"composing-text-and-image-for-image-retrieval","repo_url":"https://github.com/alinstein/Modify-image-by-text","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"composing-text-and-image-for-image-retrieval","repo_url":"https://github.com/google/tirg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"composing-text-and-image-for-image-retrieval","repo_url":"https://github.com/naver/artemis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"composing-text-and-image-for-image-retrieval","repo_url":"https://github.com/yahoo/maaf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"multi-modal","task_name":"Image Retrieval with Multi-Modal Query"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-retrieval-with-multi-modal-query-on","task":"Image Retrieval with Multi-Modal Query","dataset":"Fashion200k","model":"TIRG","rank_in_archive_order":4,"of":8,"metrics":{"Recall@1":"14.1","Recall@10":"42.5","Recall@50":"63.8"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-with-multi-modal-query-on-1","task":"Image Retrieval with Multi-Modal Query","dataset":"FashionIQ","model":"TIRG","rank_in_archive_order":2,"of":2,"metrics":{"Recall@10":"3.34"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-with-multi-modal-query-on-mit","task":"Image Retrieval with Multi-Modal Query","dataset":"MIT-States","model":"TIRG","rank_in_archive_order":2,"of":5,"metrics":{"Recall@1":"12.2","Recall@10":"43.1","Recall@5":"31.9"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1812.07119","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1812.07119"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/naver/artemis","reach":{"status":"ok","spdx":"NOASSERTION"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/alinstein/Modify-image-by-text","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google/tirg","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yahoo/maaf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":1},"by_repo_kind":{"listed":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"ca03f91d739912f2","entry":"pairwise_distances","repo":"google/tirg","repo_kind":"listed","path":"torch_functions.py","file_url":"https://github.com/google/tirg/blob/HEAD/torch_functions.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ca03f91d739912f2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}