Datasets › Test-of-Time
Test-of-Time (Test of Time Synthetic Video Dataset)
The goal of this dataset is to probe video-language models for understanding of simple temporal relations like "before" and "after". The dataset is only meant to be an evaluation set and not a training set.
Contents: 1. The dataset has synthetic videos which consists of a pair of shapes appearing gradually. For example, video for the caption "a red circle appears after a yellow circle" will first show a "yellow circle" appear and then a "red circle" appear. The model has to determine the right caption in comparison with a distractor caption "a yellow circle appears after a red circle". Note that this distractor caption has the same set of words but in a different order, motivated by the Winograd schema. 2. The dataset also has a control set in which videos only have a single event, e.g., "a red circle appears". Note that this is a control task to ensure that these videos are not out-of-distribution for a given video model. A time-aware model shall perform perfectly well on both sets. A space-aware model that is not time-aware shall perform poorly on the temporal task while performing perfectly on the control task.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Video-Text Retrieval | Test-of-Time | Video-LLAMA 2-Class Accuracy 88.33 | Video-LLaMA: An Instruction-tuned Audio-Visual Language... | damo-nlp-sg/video-llama +3 | 4 | Compare |
Papers archive 2025-07-28
4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 5. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding | 2 | 1 | 4 Dec 2023 | ran 7 of 11 samples (4 unverified) |
| Videoprompter: an ensemble of foundational models for zero-shot video understanding | 0 | 1 | 23 Oct 2023 | not harvested |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding | 4 | 1 | 5 Jun 2023 | ran 18 of 25 samples (7 unverified; 9 pointer-only for licence) |
| Test of Time: Instilling Video-Language Models with a Sense of Time | 1 | 1 | 5 Jan 2023 | ran 1 of 10 samples (9 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Test-of-Time
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections