Browse State-of-the-Art › Cross-Modal Retrieval
Cross-Modal Retrieval
244 papers with code · 13 benchmarks · 24 datasets archive 2025-07-28
Cross-Modal Retrieval (CMR) is a task of retrieving items across different modalities, such as image, text, video, and audio. The core challenge of CMR is the heterogeneity gap, which arises because data from different modalities have distinct representations, making direct comparison difficult. To address this, most CMR methods focus on learning a shared latent embedding space. In this space, concepts from different modalities are projected, allowing their similarity to be measured using a distance metric.
Scene-centric vs. Object-centric Image-Text Cross-modal Retrieval: A Reproducibility Study
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
13 leaderboard tables shown for this task, 13 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 13 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
24 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
5 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 244 papers with code (522 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
17 Jun 2019 11 repositories listed Syntology ran 2 of 29 samples · 27 unverifiedIn the second stage, SCAE predicts parameters of a few object capsules, which are then used to reconstruct part poses.
-
18 Jul 2017 10 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 4 pointer-only (licence)We present a new technique for learning visual-semantic embeddings for cross-modal retrieval.
-
23 Jun 2020 7 repositories listed Syntology ran 3 of 18 samples · 15 unverifiedThis paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS.
-
16 Jul 2021 6 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.
-
5 Feb 2021 6 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks.
-
21 Mar 2018 6 repositories listed Syntology ran 7 of 16 samples · 9 unverified · 1 pointer-only (licence)Prior work either simply aggregates the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions, or uses a multi-step attentional process to…
-
11 Feb 2021 5 repositories listed Syntology ran 8 of 10 samples · 2 unverified · 9 pointer-only (licence)In this paper, we leverage a noisy dataset of over one billion image alt-text pairs, obtained without expensive filtering or post-processing steps in the Conceptual Captions dataset.
-
12 Sep 2022 4 repositories listed Syntology ran 2 of 8 samples · 6 unverified · 2 pointer-only (licence)Although artificial intelligence (AI) has made significant progress in understanding molecules in a wide range of fields, existing models generally acquire the single cognitive ability from the single molecular modality.
-
13 Jan 2021 4 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedInstead, we propose to use Probabilistic Cross-Modal Embedding (PCME), where samples from the different modalities are represented as probabilistic distributions in the common embedding space.
-
13 Apr 2020 4 repositories listed Syntology ran 11 of 23 samples · 12 unverifiedLarge-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks.
-
1 Mar 2020 4 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedTo improve fine-grained video-text retrieval, we propose a Hierarchical Graph Reasoning (HGR) model, which decomposes video-text matching into global-to-local levels.
-
7 Dec 2014 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedOur approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences between language and visual data.
-
9 May 2023 3 repositories listed Syntology ran 24 of 34 samples · 10 unverified · 32 pointer-only (licence)We show that all combinations of paired data are not necessary to train such a joint embedding, and only image-paired data is sufficient to bind the modalities together.
-
26 Aug 2022 3 repositories listedIn this paper, we propose a simple but effective method for dealing with the challenging fine-grained cross-modal retrieval task where it aims to enable flexible retrieval among subor-dinate categories across different…
-
27 Jan 2022 3 repositories listed Syntology ran 8 of 21 samples · 13 unverifiedOur benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups.
-
3 Nov 2021 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedVision-and-language (VL) pre-training has proven to be highly effective on various VL downstream tasks.
-
31 Dec 2020 3 repositories listedExisted pre-training methods either focus on single-modal tasks or multi-modal tasks, and cannot effectively adapt to each other.
-
20 May 2020 3 repositories listedIn this paper, we address the text and image matching in cross-modal retrieval of the fashion industry.
-
19 Mar 2025 2 repositories listedHowever, progress in dermatology has lagged behind other medical domains due to the lack of standard image-text pairs.
-
29 Nov 2024 2 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)The field of computational pathology has been transformed with recent advances in foundation models that encode histopathology region-of-interests (ROIs) into versatile and transferable feature representations via…
-
30 Sep 2024 2 repositories listed Syntology ran 9 of 15 samples · 6 unverified · 15 pointer-only (licence)Surgical video-language pretraining (VLP) faces unique challenges due to the knowledge domain gap and the scarcity of multi-modal data.
-
10 Jun 2024 2 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedHowever, current medical VLMs are generally limited to 2D images and short reports, and do not leverage electronic health record (EHR) data for supervision.
-
26 Dec 2023 2 repositories listedIn this work, we present LeanVec, a framework that combines linear dimensionality reduction with vector quantization to accelerate similarity search on high-dimensional vectors while maintaining accuracy.
-
20 Jun 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedFrom YouTube, we curate QUILT: a large-scale vision-language dataset consisting of $802, 144$ image and text pairs.
-
12 Jun 2023 2 repositories listedFurthermore, as the diversity and differentiation of remote sensing scenes weaken the understanding of scenes, a new metric, namely, scene recall is proposed to measure the perception of scenes by evaluating scene-level…
-
6 Jun 2023 2 repositories listedIn this study, we introduce MolFM, a multimodal molecular foundation model designed to facilitate joint representation learning from molecular structures, biomedical texts, and knowledge graphs.
-
29 May 2023 2 repositories listed Syntology ran 15 of 42 samples · 27 unverifiedBased on the proposed VAST-27M dataset, we train an omni-modality video-text foundational model named VAST, which can perceive and process vision, audio, and subtitle modalities from video, and better support various…
-
4 Apr 2023 2 repositories listedThis paper presents the AToMiC (Authoring Tools for Multimedia Content) dataset, designed to advance research in image/text cross-modal retrieval.
-
6 Dec 2022 2 repositories listedThe rich semantics are further regarded as semantic prior to trigger the learning of Diffusion Transformer, which produces the output sentence in a diffusion process.
-
6 Dec 2022 2 repositories listedTo verify the effectiveness of our approach, extensive experiments are conducted on MS-COCO, CUB Captions, and Flickr30K, which are commonly used in cross-modal retrieval.
Syntology lines on 20 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections