Datasets › MSR-VTT › Papers where code ran, page 1

MSR-VTT

Papers archive 2025-07-28

papers with a benchmark row: 117 · with a code link: 92 · where Syntology ran a sample: 49 (42 with a run with no instrument failure, 7 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (49 of 117 with a benchmark row: 42 with a run with no instrument failure, 7 where every run was a failure of Syntology's instrument)

Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.

Page 1 of 1: papers 1 to 49 of the 49 papers with a benchmark row here where Syntology ran at least one harvested sample (42 with a run with no instrument failure, 7 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 640. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
Make Your Training Flexible: Towards Deployment-Efficient Video Models 1 2 18 Mar 2025 official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 0 violated, 17 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (6 pointer-only for licence)
Gramian Multimodal Representation Learning and Alignment 2 2 16 Dec 2024 official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 1 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer 1 2 20 Mar 2024 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (5 pointer-only for licence)
Multi-granularity Correspondence Learning from Long-term Noisy Videos 1 1 30 Jan 2024 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 2 samples that ran constructed an object rather than computing a result
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing 1 1 29 Nov 2023 official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 2 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (12 pointer-only for licence)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale 1 3 7 Oct 2023 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (5 pointer-only for licence)
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment 6 2 3 Oct 2023 official (archive's flag): 5 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (3 pointer-only for licence)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval 1 1 29 Sep 2023 official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 1 violated, 7 with no contract checked; 7 where Syntology's instrument failed) · 2 unverified (8 pointer-only for licence)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 1 27 Sep 2023 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation 1 1 27 Sep 2023 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
Accurate and Fast Compressed Video Captioning 1 1 22 Sep 2023 official (archive's flag): 10 ran · 10 ran (of which 8 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified
Unified Coarse-to-Fine Alignment for Video-Text Retrieval 1 2 18 Sep 2023 official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (4 pointer-only for licence)
ModelScope Text-to-Video Technical Report 5 1 12 Aug 2023 official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence)
VideoComposer: Compositional Video Synthesis with Motion Controllability 4 1 3 Jun 2023 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence)
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 3 29 May 2023 official (archive's flag): 12 ran · 35 ran (of which 4 constructed an object rather than computing a result; 29 with no instrument failure: 2 honoured, 1 violated, 26 with no contract checked; 6 where Syntology's instrument failed) · 7 unverified (8 pointer-only for licence)
ImageBind: One Embedding Space To Bind Them All 3 1 9 May 2023 official (archive's flag): 18 ran · 24 ran (of which 14 constructed an object rather than computing a result; 21 with no instrument failure: 1 honoured, 1 violated, 19 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (32 pointer-only for licence)
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models 4 2 18 Apr 2023 official: no sample here; runs from other or unrecorded repositories · 18 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 2 honoured, 3 violated, 7 with no contract checked; 6 where Syntology's instrument failed) · 8 unverified (3 pointer-only for licence)
MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks 1 1 29 Mar 2023 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models 1 2 28 Mar 2023 official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence)
Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning 4 1 25 Mar 2023 official (archive's flag): 1 ran · 12 ran (of which 7 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified
MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models 1 7 23 Mar 2023 official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result
DiffusionRet: Generative Text-Video Retrieval with Diffusion Model 4 2 17 Mar 2023 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 4 3 1 Feb 2023 community repositories only · 17 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence)
Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval? 4 1 31 Dec 2022 official (archive's flag): 14 ran · 20 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 2 honoured, 1 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 5 unverified (11 pointer-only for licence)
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 2 6 Dec 2022 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified
X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks 2 2 22 Nov 2022 official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (6 pointer-only for licence)
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations 4 3 21 Nov 2022 official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified
CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment 1 1 14 Sep 2022 official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence)
TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval 1 1 16 Jul 2022 official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (4 pointer-only for licence)
X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval 3 1 15 Jul 2022 official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence)
Revealing Single Frame Bias for Video-and-Language Learning 2 3 7 Jun 2022 community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
GIT: A Generative Image-to-text Transformer for Vision and Language 1 1 27 May 2022 official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 1 4 May 2022 10 ran (of which 5 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (2 pointer-only for licence)
X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval 1 1 28 Mar 2022 official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (2 pointer-only for licence)
Disentangled Representation Learning for Text-Video Retrieval 2 1 14 Mar 2022 official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (4 pointer-only for licence)
Bridging Video-text Retrieval with Multiple Choice Questions 2 3 13 Jan 2022 official (archive's flag): 2 ran · 13 ran (of which 8 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 11 unverified (6 pointer-only for licence)
Cross Modal Retrieval with Querybank Normalisation 1 1 23 Dec 2021 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion 1 1 24 Nov 2021 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text 5 1 22 Apr 2021 community repositories only · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (8 pointer-only for licence)
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval 5 4 18 Apr 2021 official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence)
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval 5 3 1 Apr 2021 official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence)
MDMMT: Multidomain Multimodal Transformer for Video Retrieval 3 2 19 Mar 2021 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (5 pointer-only for licence)
Multi-modal Transformer for Video Retrieval 1 3 21 Jul 2020 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation 2 1 15 Feb 2020 official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (1 pointer-only for licence)
End-to-End Learning of Visual Representations from Uncurated Instructional Videos 4 1 13 Dec 2019 official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
Use What You Have: Video Retrieval Using Representations From Collaborative Experts 3 2 31 Jul 2019 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence)
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips 4 3 7 Jun 2019 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
A Joint Sequence Fusion Model for Video Question Answering and Retrieval 2 2 7 Aug 2018 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified