Browse State-of-the-Art › Text to Video Retrieval
Text to Video Retrieval
51 papers with code · 3 benchmarks · 7 datasets archive 2025-07-28
She's gone I can't find her anywhere I'm looking everywhere for her Everywhere is dark
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Kinetics-GEB+ (2 rows) | FROZEN-revised | GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding... | code | — | Compare |
| MSR-VTT (1 row) | CLIP4Clip | CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval | code | Syntology ran 3 of 4 samples · 1 unverified | Compare |
| MSVD-Indonesian (1 row) | X-CLIP (Cross-Lingual) | MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 51 papers with code (75 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Apr 2021 5 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and…
-
18 Apr 2021 5 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 3 pointer-only (licence)In this paper, we propose a CLIP4Clip model to transfer the knowledge of the CLIP model to video-language retrieval in an end-to-end manner.
-
1 Apr 2021 5 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedOur objective in this work is video-text retrieval - in particular a joint embedding that enables efficient text-to-video retrieval.
-
13 Dec 2019 4 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedAnnotating videos is cumbersome, expensive and not scalable.
-
7 Jun 2019 4 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this work, we propose instead to learn such embeddings from video data with readily available natural language annotations in the form of automatically transcribed narrations.
-
19 Mar 2021 3 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)We present a new state-of-the-art on the text to video retrieval task on MSRVTT and LSMDC benchmarks where our model outperforms all previous solutions by a large margin.
-
29 Oct 2024 2 repositories listedWe show that our system, without joint training, achieves better or comparable results to state-of-the-art models and commercial solutions on multiple text-to-video retrieval benchmarks.
-
22 Nov 2022 2 repositories listed Syntology ran 2 of 6 samples · 4 unverified · 6 pointer-only (licence)Vision language pre-training aims to learn alignments between vision and language from a large amount of data.
-
7 Jun 2022 2 repositories listed Syntology ran 4 of 12 samples · 8 unverifiedTraining an effective video-and-language model intuitively requires multiple frames as model inputs.
-
24 Mar 2022 2 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedLarge-scale pretrained image-text models have shown incredible zero-shot performance in a handful of tasks, including video ones such as action recognition and text-to-video retrieval.
-
15 Mar 2022 2 repositories listedRecent dominant methods for video-language pre-training (VLP) learn transferable representations from the raw pixels in an end-to-end manner to achieve advanced performance on downstream video-language retrieval.
-
13 Jan 2022 2 repositories listed Syntology ran 13 of 24 samples · 11 unverified · 6 pointer-only (licence)As an additional benefit, our method achieves competitive results with much shorter pre-training videos on single-modality downstream tasks, e.
-
15 Apr 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Partially relevant video retrieval (PRVR) is a practical yet challenging task in text-to-video retrieval, where videos are untrimmed and contain much background content.
-
7 Apr 2025 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 1 pointer-only (licence)To filter unnecessary similarity interactions and decrease trainable parameters in the Interactive Similarity Aggregation (ISA) module, we design a Similarity Reorganization (SR) module to identify attentive…
-
13 Mar 2025 1 repository listedText-to-Video Retrieval (TVR) aims to match videos with corresponding textual queries, yet the continual influx of new video content poses a significant challenge for maintaining system performance over time.
-
1 Jan 2024 1 repository listedTo address this issue, we adopt multi-granularity visual feature learning, ensuring the model's comprehensiveness in capturing visual content features spanning from abstract to detailed levels during the training phase.
-
1 Jan 2024 1 repository listedFor text-to-video retrieval (T2VR) which aims to retrieve unlabeled videos by ad-hoc textual queries CLIP-based methods currently lead the way.
-
15 Nov 2023 1 repository listed Syntology ran 1 of 3 samples · 2 unverifiedDespite being (pre)trained on a massive amount of data, state-of-the-art video-language alignment models are not robust to semantically-plausible contrastive changes in the video captions.
-
8 Oct 2023 1 repository listedDespite significant results achieved by Contrastive Language-Image Pretraining (CLIP) in zero-shot image recognition, limited effort has been made exploring its potential for zero-shot video recognition.
-
29 Sep 2023 1 repository listed Syntology ran 12 of 19 samples · 7 unverifiedIn this paper, we propose a novel Prototype-based Aleatoric Uncertainty Quantification (PAU) framework to provide trustworthy predictions by quantifying the uncertainty arisen from the inherent data ambiguity.
-
18 Sep 2023 1 repository listed Syntology ran 8 of 14 samples · 6 unverifiedSpecifically, our model captures the cross-modal similarity information at different granularity levels.
-
20 Jun 2023 1 repository listedSince the availability of the pretraining resources with Indonesian sentences is relatively limited, the applicability of those approaches to our dataset is still questionable.
-
23 Mar 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedTherefore, we propose MEta Loss TRansformer (MELTR), a plug-in module that automatically and non-linearly combines various loss functions to aid learning the target task via auxiliary learning.
-
4 Feb 2023 1 repository listedThis paper presents a new method for end-to-end Video Question Answering (VideoQA), aside from the current popularity of using large-scale pre-training with huge feature extractors.
-
1 Jan 2023 1 repository listedDuring the knowledge distillation, an inheritance student branch is devised to absorb the knowledge from the teacher model.
-
9 Dec 2022 1 repository listedFurthermore, our model also obtains state-of-the-art video question-answering results on ActivityNet-QA, MSRVTT-QA, MSRVTT-MC and TVQA.
-
21 Nov 2022 1 repository listedIn this paper we tackle the cross-modal video retrieval problem and, more specifically, we focus on text-to-video retrieval.
-
4 Sep 2022 1 repository listedMasked visual modeling (MVM) has been recently proven effective for visual pre-training.
-
26 Aug 2022 1 repository listedTo fill the gap, we propose in this paper a novel T2VR subtask termed Partially Relevant Video Retrieval (PRVR).
-
16 Jul 2022 1 repository listedWe then introduce \textbf{Clover}\textemdash a Correlated Video-Language pre-training method\textemdash towards a universal Video-Language model for solving multiple video understanding tasks with neither performance…
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections