Datasets › MSR-VTT

MSR-VTT

Introduced by Jun Xu et al. in MSR-VTT: A Large Video Description Dataset for Bridging Video and Language1 Jan 2016 archive 2025-07-28

MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon Mechanical Turks. There are about 29,000 unique words in all captions. The standard splits uses 6,513 clips for training, 497 clips for validation, and 2,990 clips for testing.

Source: Learning to Discretely Compose Reasoning Module Networksfor Video Captioning

Benchmarks archive 2025-07-28

All 8 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 117 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 640. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Make Your Training Flexible: Towards Deployment-Efficient Video Models 1 2 18 Mar 2025 ran 7 of 17 samples (10 unverified)
Gramian Multimodal Representation Learning and Alignment 2 2 16 Dec 2024 ran 2 of 12 samples (10 unverified)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs 1 1 11 Apr 2024 not harvested
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 3 22 Mar 2024 not harvested
vid-TLDR: Training Free Token merging for Light-weight Video Transformer 1 2 20 Mar 2024 ran 3 of 4 samples (1 unverified)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis 0 2 22 Feb 2024 not harvested
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Multi-granularity Correspondence Learning from Long-term Noisy Videos 1 1 30 Jan 2024 ran 2 of 3 samples (1 unverified)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning 1 1 1 Jan 2024 not harvested
Holistic Features are almost Sufficient for Text-to-Video Retrieval 1 2 1 Jan 2024 not harvested
A Recipe for Scaling up Text-to-Video Generation with Text-free Videos 1 1 25 Dec 2023 not harvested
VideoPoet: A Large Language Model for Zero-Shot Video Generation 0 1 21 Dec 2023 not harvested
Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation 1 1 7 Dec 2023 not harvested
RTQ: Rethinking Video-language Understanding Based on Image-text Model 2 2 1 Dec 2023 not harvested
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing 1 1 29 Nov 2023 ran 10 of 12 samples (2 unverified; 12 pointer-only for licence)
Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning 2 1 27 Nov 2023 not harvested
Make Pixels Dance: High-Dynamic Video Generation 0 1 18 Nov 2023 not harvested
OmniVec: Learning robust representations with cross modal sharing 0 2 7 Nov 2023 not harvested
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale 1 3 7 Oct 2023 ran 4 of 5 samples (1 unverified; 5 pointer-only for licence)
IcoCap: Improving Video Captioning by Compounding Images 0 2 5 Oct 2023 not harvested
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment 6 2 3 Oct 2023 ran 7 of 14 samples (7 unverified)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval 1 1 29 Sep 2023 ran 12 of 19 samples (7 unverified)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation 1 1 27 Sep 2023 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 1 27 Sep 2023 ran 1 of 4 samples (3 unverified)
Accurate and Fast Compressed Video Captioning 1 1 22 Sep 2023 ran 10 of 14 samples (4 unverified)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning 1 1 20 Sep 2023 not harvested
Unified Coarse-to-Fine Alignment for Video-Text Retrieval 1 2 18 Sep 2023 ran 8 of 14 samples (6 unverified)
ModelScope Text-to-Video Technical Report 5 1 12 Aug 2023 ran 11 of 16 samples (5 unverified)
Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment 0 1 24 Jul 2023 not harvested
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model 1 2 15 Jun 2023 not harvested

The full list of 117 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MSR-VTT
  • MSR-VTT-1kA
  • MSRVTT

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections