Datasets › MSR-VTT › Papers, page 2

MSR-VTT

Papers archive 2025-07-28

papers with a benchmark row: 117 · with a code link: 92 · where Syntology ran a sample: 49 (42 with a run with no instrument failure, 7 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (49 of 117 with a benchmark row: 42 with a run with no instrument failure, 7 where every run was a failure of Syntology's instrument)

Page 2 of 2: papers 101 to 117 of 117 with a leaderboard row on this dataset's benchmarks, newest first by the archive's date (ties by slug; undated papers last).

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 640. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding 1 1 20 May 2021 not harvested
GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions 1 1 30 Apr 2021 not harvested
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text 5 1 22 Apr 2021 community repositories only · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (8 pointer-only for licence)
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval 5 4 18 Apr 2021 official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence)
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval 5 3 1 Apr 2021 official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence)
MDMMT: Multidomain Multimodal Transformer for Video Retrieval 3 2 19 Mar 2021 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (5 pointer-only for licence)
A Straightforward Framework For Video Retrieval Using CLIP 1 2 24 Feb 2021 not harvested
Multi-modal Transformer for Video Retrieval 1 3 21 Jul 2020 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Noise Estimation Using Density Estimation for Self-Supervised Multimodal Learning 1 1 6 Mar 2020 not harvested
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation 2 1 15 Feb 2020 official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (1 pointer-only for licence)
End-to-End Learning of Visual Representations from Uncurated Instructional Videos 4 1 13 Dec 2019 official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
Use What You Have: Video Retrieval Using Representations From Collaborative Experts 3 2 31 Jul 2019 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence)
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips 4 3 7 Jun 2019 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
A Joint Sequence Fusion Model for Video Question Answering and Retrieval 2 2 7 Aug 2018 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text Retrieval 1 1 11 Jun 2018 not harvested
Temporal Tessellation: A Unified Approach for Video Analysis 1 1 21 Dec 2016 not harvested
Learning Language-Visual Embedding for Movie Understanding with Natural-Language 0 1 26 Sep 2016 not harvested