Datasets › Vinoground

Vinoground

3 Oct 2024 archive 2025-07-28

A temporal counterfactual dataset composing of 1000 short and natural video-caption pairs.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Temporal Relation Extraction Vinoground GPT-4o (CoT) Text Score 59.2 — — 24 Compare

Papers archive 2025-07-28

12 shown of 12 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 17. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 8 2 18 Sep 2024 ran 8 of 12 samples (4 unverified)
LLaVA-OneVision: Easy Visual Task Transfer 2 2 6 Aug 2024 not harvested
MiniCPM-V: A GPT-4V Level MLLM on Your Phone 2 1 3 Aug 2024 ran 9 of 14 samples (5 unverified)
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 1 2 3 Jul 2024 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 3 1 11 Jun 2024 ran 7 of 17 samples (10 unverified)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 1 1 8 Apr 2024 ran 7 of 9 samples (2 unverified)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 1 2 8 Mar 2024 not harvested
VTimeLLM: Empower LLM to Grasp Video Moments 1 1 30 Nov 2023 ran 5 of 11 samples (6 unverified; 11 pointer-only for licence)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 6 1 16 Nov 2023 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment 6 1 3 Oct 2023 ran 7 of 14 samples (7 unverified)
ImageBind: One Embedding Space To Bind Them All 3 1 9 May 2023 ran 24 of 34 samples (10 unverified; 32 pointer-only for licence)
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding 2 1 28 Sep 2021 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Vinoground

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections