{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/compositional-learning-of-image-text-query","title":"Compositional Learning of Image-Text Query for Image Retrieval","arxiv_id":"2006.11149","date":"2020-06-19","proceeding":null,"authors":["Muhammad Umer Anwaar","Egor Labintcev","Martin Kleinsteuber"],"abstract":"In this paper, we investigate the problem of retrieving images from a database based on a multi-modal (image-text) query. Specifically, the query text prompts some modification in the query image and the task is to retrieve images with the desired modifications. For instance, a user of an E-Commerce platform is interested in buying a dress, which should look similar to her friend's dress, but the dress should be of white color with a ribbon sash. In this case, we would like the algorithm to retrieve some dresses with desired modifications in the query dress. We propose an autoencoder based model, ComposeAE, to learn the composition of image and text query for retrieving images. We adopt a deep metric learning approach and learn a metric that pushes composition of source image and text query closer to the target images. We also propose a rotational symmetry constraint on the optimization problem. Our approach is able to outperform the state-of-the-art method TIRG \\cite{TIRG} on three benchmark datasets, namely: MIT-States, Fashion200k and Fashion IQ. In order to ensure fair comparison, we introduce strong baselines by enhancing TIRG method. To ensure reproducibility of the results, we publish our code here: \\url{https://github.com/ecom-research/ComposeAE}.","url_abs":"https://arxiv.org/abs/2006.11149v3","url_pdf":"https://arxiv.org/pdf/2006.11149v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"compositional-learning-of-image-text-query","repo_url":"https://github.com/ecom-research/ComposeAE","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"multi-modal","task_name":"Image Retrieval with Multi-Modal Query"},{"task_slug":"metric-learning","task_name":"Metric Learning"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-retrieval-on-fashion-iq","task":"Image Retrieval","dataset":"Fashion IQ","model":"ComposeAE","rank_in_archive_order":22,"of":22,"metrics":{"(Recall@10+Recall@50)/2":"20.6"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-with-multi-modal-query-on","task":"Image Retrieval with Multi-Modal Query","dataset":"Fashion200k","model":"ComposeAE","rank_in_archive_order":2,"of":8,"metrics":{"Recall@1":"22.8","Recall@10":"55.3","Recall@50":"73.4"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-with-multi-modal-query-on-1","task":"Image Retrieval with Multi-Modal Query","dataset":"FashionIQ","model":"ComposeAE","rank_in_archive_order":1,"of":2,"metrics":{"Recall@10":"11.8"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-with-multi-modal-query-on-mit","task":"Image Retrieval with Multi-Modal Query","dataset":"MIT-States","model":"ComposeAE","rank_in_archive_order":1,"of":5,"metrics":{"Recall@1":"13.9","Recall@10":"47.9","Recall@5":"35.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2006.11149","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2006.11149"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ecom-research/ComposeAE","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"ca03f91d739912f2","entry":"pairwise_distances","repo":"ecom-research/ComposeAE","repo_kind":"official","path":"torch_functions.py","file_url":"https://github.com/ecom-research/ComposeAE/blob/HEAD/torch_functions.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ca03f91d739912f2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}