Browse State-of-the-Art › Image Retrieval with Multi-Modal Query
Image Retrieval with Multi-Modal Query
9 papers with code · 3 benchmarks · 2 datasets archive 2025-07-28
The problem of retrieving images from a database based on a multi-modal (image- text) query. Specifically, the query text prompts some modification in the query image and the task is to retrieve images with the desired modifications.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Fashion200k (8 rows) | Css-Net | Collaborative Group: Composed Image Retrieval via Consensus... | — | — | Compare |
| MIT-States (5 rows) | ComposeAE | Compositional Learning of Image-Text Query for Image Retrieval | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
| FashionIQ (2 rows) | ComposeAE | Compositional Learning of Image-Text Query for Image Retrieval | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
9 shown of 9 papers with code (10 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
17 Nov 2014 74 repositories listed Syntology ran 13 of 34 samples · 21 unverified · 6 pointer-only (licence)Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions.
-
5 Jun 2017 20 repositories listed Syntology ran 3 of 7 samples · 4 unverified · 1 pointer-only (licence)Relational reasoning is a central component of generally intelligent behavior, but has proven difficult for neural networks to learn.
-
22 Sep 2017 7 repositories listed Syntology ran 9 of 10 samples · 1 unverified · 7 pointer-only (licence)We introduce a general-purpose conditioning method for neural networks called FiLM: Feature-wise Linear Modulation.
-
18 Dec 2018 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedIn this paper, we study the task of image retrieval, where the input query is specified in the form of an image plus some text that describes desired modifications to the input image.
-
14 Nov 2022 1 repository listed Syntology ran 3 of 10 samples · 7 unverifiedThe key idea underpinning the proposed method is to integrate fine- and coarse-grained retrieval as matching data points with small and large fluctuations, respectively.
-
19 Jun 2020 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedIn this paper, we investigate the problem of retrieving images from a database based on a multi-modal (image-text) query.
-
27 Mar 2018 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedIn addition, we show that not only can our model recognize unseen compositions robustly in an open-world setting, it can also generalize to compositions where objects themselves were unseen during training.
-
3 Aug 2017 1 repository listedThis paper proposes an automatic spatially-aware concept discovery approach using weakly labeled image-text data from shopping websites.
-
18 Nov 2015 1 repository listedWe tackle image question answering (ImageQA) problem by learning a convolutional neural network (CNN) with a dynamic parameter layer whose weights are determined adaptively based on questions.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections