Datasets › Test-of-Time

Test-of-Time (Test of Time Synthetic Video Dataset)

Introduced by Piyush Bagad et al. in Test of Time: Instilling Video-Language Models with a Sense of Time5 Jan 2023 archive 2025-07-28

The goal of this dataset is to probe video-language models for understanding of simple temporal relations like "before" and "after". The dataset is only meant to be an evaluation set and not a training set.

Contents: 1. The dataset has synthetic videos which consists of a pair of shapes appearing gradually. For example, video for the caption "a red circle appears after a yellow circle" will first show a "yellow circle" appear and then a "red circle" appear. The model has to determine the right caption in comparison with a distractor caption "a yellow circle appears after a red circle". Note that this distractor caption has the same set of words but in a different order, motivated by the Winograd schema. 2. The dataset also has a control set in which videos only have a single event, e.g., "a red circle appears". Note that this is a control task to ensure that these videos are not out-of-distribution for a given video model. A time-aware model shall perform perfectly well on both sets. A space-aware model that is not time-aware shall perform poorly on the temporal task while performing perfectly on the control task.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video-Text Retrieval Test-of-Time Video-LLAMA 2-Class Accuracy 88.33 Video-LLaMA: An Instruction-tuned Audio-Visual Language... damo-nlp-sg/video-llama +3 4 Compare

Papers archive 2025-07-28

4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 5. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding 2 1 4 Dec 2023 ran 7 of 11 samples (4 unverified)
Videoprompter: an ensemble of foundational models for zero-shot video understanding 0 1 23 Oct 2023 not harvested
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding 4 1 5 Jun 2023 ran 18 of 25 samples (7 unverified; 9 pointer-only for licence)
Test of Time: Instilling Video-Language Models with a Sense of Time 1 1 5 Jan 2023 ran 1 of 10 samples (9 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Test-of-Time

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections