Datasets › VATEX

VATEX (Video And TEXt)

Introduced by Xin Wang et al. in VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research1 Jan 2019 archive 2025-07-28

VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions. It has two tasks for video-and-language research: (1) Multilingual Video Captioning, aimed at describing a video in various languages with a compact unified captioning model, and (2) Video-guided Machine Translation, to translate a source language description into the target language using the video information as additional spatiotemporal context.

Source: https://arxiv.org/pdf/1904.03493.pdf Image Source: https://arxiv.org/pdf/1904.03493.pdf

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

20 shown of 20 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 118. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Gramian Multimodal Representation Learning and Alignment 2 2 16 Dec 2024 ran 2 of 12 samples (10 unverified)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 3 22 Mar 2024 not harvested
Holistic Features are almost Sufficient for Text-to-Video Retrieval 1 1 1 Jan 2024 not harvested
Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning 2 1 27 Nov 2023 not harvested
IcoCap: Improving Video Captioning by Compounding Images 0 2 5 Oct 2023 not harvested
Accurate and Fast Compressed Video Captioning 1 1 22 Sep 2023 ran 10 of 14 samples (4 unverified)
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model 1 1 15 Jun 2023 not harvested
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 2 29 May 2023 ran 15 of 42 samples (27 unverified)
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset 1 2 17 Apr 2023 not harvested
Unmasked Teacher: Towards Training-Efficient Video Foundation Models 1 1 28 Mar 2023 ran 3 of 8 samples (5 unverified)
Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval? 4 1 31 Dec 2022 ran 14 of 25 samples (11 unverified)
VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners 0 2 9 Dec 2022 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 2 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Diverse Video Captioning by Adaptive Spatio-temporal Attention 1 1 19 Aug 2022 not harvested
TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval 1 1 16 Jul 2022 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
Cross Modal Retrieval with Querybank Normalisation 1 1 23 Dec 2021 ran 2 of 3 samples (1 unverified)
Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval 1 1 3 Dec 2021 not harvested
CLIP2Video: Mastering Video-Text Retrieval via Image CLIP 1 1 21 Jun 2021 ran 0 of 1 samples (1 unverified)
NITS-VC System for VATEX Video Captioning Challenge 2020 0 1 7 Jun 2020 not harvested
Object Relational Graph with Teacher-Recommended Learning for Video Captioning 0 1 26 Feb 2020 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • VATEX

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections