Browse State-of-the-Art › Video Retrieval
Video Retrieval
255 papers with code · 19 benchmarks · 35 datasets archive 2025-07-28
The objective of video retrieval is as follows: given a text query and a pool of candidate videos, select the video which corresponds to the text query. Typically, the videos are returned as a ranked list of candidates and scored via document retrieval metrics.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
19 leaderboard tables shown for this task, 19 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 19 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
35 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 35 until expanded.
Subtasks archive 2025-07-28
5 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 255 papers with code (486 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
4 May 2022 6 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedWe apply a contrastive loss between unimodal image and text embeddings, in addition to a captioning loss on the multimodal decoder outputs which predicts text tokens autoregressively.
-
24 Apr 2018 6 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIn this paper, we introduce a network architecture that takes long-term content into account and enables fast per-video processing at the same time.
-
22 Apr 2021 5 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and…
-
18 Apr 2021 5 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 3 pointer-only (licence)In this paper, we propose a CLIP4Clip model to transfer the knowledge of the CLIP model to video-language retrieval in an end-to-end manner.
-
1 Apr 2021 5 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedOur objective in this work is video-text retrieval - in particular a joint embedding that enables efficient text-to-video retrieval.
-
7 Apr 2018 5 repositories listedWe evaluate our method on the task of video retrieval and report results for the MPII Movie Description and MSR-VTT datasets.
-
20 May 2023 4 repositories listed Syntology ran 9 of 23 samples · 14 unverifiedIn this paper, we propose the Disentangled Conceptualization and Set-to-set Alignment (DiCoSA) to simulate the conceptualizing and reasoning process of human beings.
-
Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning25 Mar 2023 4 repositories listed Syntology ran 11 of 16 samples · 5 unverifiedContrastive learning-based video-language representation learning approaches, e.
-
17 Mar 2023 4 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedExisting text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.
-
1 Feb 2023 4 repositories listed Syntology ran 9 of 19 samples · 10 unverifiedIn contrast to predominant paradigms of solely relying on sequence-to-sequence generation or encoder-based instance discrimination, mPLUG-2 introduces a multi-module composition network by sharing common universal…
-
31 Dec 2022 4 repositories listed Syntology ran 14 of 25 samples · 11 unverifiedMost existing text-video retrieval methods focus on cross-modal matching between the visual content of videos and textual query sentences.
-
21 Nov 2022 4 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedMost video-and-language representation learning approaches employ contrastive learning, e.
-
13 Dec 2019 4 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedAnnotating videos is cumbersome, expensive and not scalable.
-
7 Jun 2019 4 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this work, we propose instead to learn such embeddings from video data with readily available natural language annotations in the form of automatically transcribed narrations.
-
2 May 2017 4 repositories listedWe also introduce ActivityNet Captions, a large-scale benchmark for dense-captioning events.
-
28 Sep 2022 3 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedIn this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representation learning with minimal…
-
15 Jul 2022 3 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 1 pointer-only (licence)However, cross-grained contrast, which is the contrast between coarse-grained representations and fine-grained representations, has rarely been explored in prior research.
-
19 Mar 2021 3 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)We present a new state-of-the-art on the text to video retrieval task on MSRVTT and LSMDC benchmarks where our model outperforms all previous solutions by a large margin.
-
18 Mar 2021 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Current video retrieval efforts all found their evaluation on an instance-based assumption, that only a single caption is relevant to a query video and vice versa.
-
1 May 2020 3 repositories listed Syntology ran 5 of 13 samples · 8 unverified · 8 pointer-only (licence)We present HERO, a novel framework for large-scale video+language omni-representation learning.
-
31 Jul 2019 3 repositories listed Syntology ran 0 of 5 samples · 5 unverified · 1 pointer-only (licence)The rapid growth of video on the internet has made searching for video content using natural language queries a significant challenge.
-
28 Mar 2019 3 repositories listedIn this work, we study robust deep learning against abnormal training data from the perspective of example weighting built in empirical loss functions, i.
-
16 Dec 2024 2 repositories listed Syntology ran 2 of 12 samples · 10 unverifiedThese models typically align each modality to a designated anchor without ensuring the alignment of all modalities with each other, leading to suboptimal performance in tasks requiring a joint understanding of multiple…
-
29 Oct 2024 2 repositories listedWe show that our system, without joint training, achieves better or comparable results to state-of-the-art models and commercial solutions on multiple text-to-video retrieval benchmarks.
-
22 Mar 2024 2 repositories listedWe introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue.
-
19 Jan 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)(2) Equipping the visual and text encoder with separated prompts failed to mitigate the visual-text modality gap.
-
1 Dec 2023 2 repositories listedRecent advancements in video-language understanding have been established on the foundation of image-text models, resulting in promising outcomes due to the shared knowledge between images and videos.
-
27 Nov 2023 2 repositories listedIn this paper, we present a novel Spatial-Temporal Side Network for memory-efficient fine-tuning large image models to video understanding, named Side4Video.
-
27 Jul 2023 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning.
-
28 Jun 2023 2 repositories listedThe study is performed on two categories of video retrieval models: (i) which are pre-trained on video-text pairs and fine-tuned on downstream video retrieval datasets (Eg.
Syntology lines on 22 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections