Browse State-of-the-Art › Image-text Retrieval
Image-text Retrieval
131 papers with code · 0 benchmarks · 5 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 131 papers with code (248 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Jan 2022 9 repositories listedFurthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision.
-
25 Sep 2019 7 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.
-
15 Dec 2022 6 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedVision Transformers convert images to sequences by slicing them into patches.
-
16 Jul 2021 6 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.
-
11 Feb 2021 5 repositories listed Syntology ran 8 of 10 samples · 2 unverified · 9 pointer-only (licence)In this paper, we leverage a noisy dataset of over one billion image alt-text pairs, obtained without expensive filtering or post-processing steps in the Conceptual Captions dataset.
-
24 May 2022 3 repositories listedLarge-scale pretrained foundation models have been an emerging paradigm for building artificial intelligence (AI) systems, which can be quickly adapted to a wide range of downstream tasks.
-
2 Mar 2021 3 repositories listedFirst, WIT is the largest multimodal dataset by the number of image-text examples by 3x (at the time of writing).
-
13 Jan 2025 2 repositories listed Syntology ran 0 of 20 samples · 20 unverifiedThe development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets.
-
11 Jun 2024 2 repositories listed Syntology ran 7 of 14 samples · 7 unverifiedContrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from websites.
-
31 Mar 2024 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedAdditionally, we propose M3D-LaMed, a versatile multi-modal large language model for 3D medical image analysis.
-
19 Oct 2023 2 repositories listed Syntology ran 8 of 16 samples · 8 unverifiedThis paper reveals that large language models (LLMs), despite being trained solely on textual data, are surprisingly strong encoders for purely visual tasks in the absence of language.
-
10 Oct 2023 2 repositories listedBased on flexible spatial awareness, we further propose the Self-Expanding triplet Loss (SEL) to expand the representation space of samples and optimize the alignment of embedding space.
-
15 Jun 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)However, the compositional reasoning abilities of existing VLMs remains subpar.
-
18 May 2023 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedIn this work, we explore a scalable way for building a general representation model toward unlimited modalities.
-
11 May 2023 2 repositories listed Syntology ran 5 of 8 samples · 3 unverifiedWe present Region-aware Open-vocabulary Vision Transformers (RO-ViT) - a contrastive image-text pretraining recipe to bridge the gap between image-level pretraining and open-vocabulary object detection.
-
18 Apr 2023 2 repositories listed Syntology ran 6 of 15 samples · 9 unverified · 14 pointer-only (licence)Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept "dog" entails all images that contain dogs.
-
4 Apr 2023 2 repositories listedThis paper presents the AToMiC (Authoring Tools for Multimedia Content) dataset, designed to advance research in image/text cross-modal retrieval.
-
13 Mar 2023 2 repositories listed Syntology ran 6 of 13 samples · 7 unverifiedFoundation models trained on large-scale dataset gain a recent surge in CV and NLP.
-
13 Feb 2023 2 repositories listed Syntology ran 8 of 10 samples · 2 unverified · 10 pointer-only (licence)Particularly, on the MSRVTT retrieval task, UniAdapter achieves 49.
-
31 Jan 2023 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedReal-world data contains a vast amount of multimodal information, among which vision and language are the two most representative modalities.
-
11 Jan 2023 2 repositories listedIt is not easy to propose a new model with a novel architecture and intensively train it on a massive dataset with many GPUs to surpass many SOTA models, which are already available to use on the Internet.
-
21 Oct 2022 2 repositories listedIn the event that the gradients are not integrable to a valid loss function, we implement our proposed objectives such that they would directly operate in the gradient space instead of on the losses in the embedding…
-
3 Nov 2021 2 repositories listedWe present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network.
-
16 Mar 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedMultimodal pre-training has propelled great advancement in vision-and-language research.
-
1 Jan 2021 2 repositories listedIn recent years, the growing number of medical imaging studies is placing an ever-increasing burden on radiologists.
-
11 Jun 2020 2 repositories listed Syntology ran 10 of 20 samples · 10 unverified · 6 pointer-only (licence)We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning.
-
26 Apr 2019 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedTo bridge the learning of two modules, we use a neuro-symbolic reasoning module that executes these programs on the latent scene representation.
-
11 Jun 2025 1 repository listedDual encoder Vision-Language Models (VLM) such as CLIP are widely used for image-text retrieval tasks.
-
10 Jun 2025 1 repository listedWe present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question…
-
3 Jun 2025 1 repository listedThis paper studies the vulnerabilities of vision foundation models, focusing specifically on CLIP and ViTs, and explores the transferability of adversarial attacks to downstream tasks.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections