{"url":"/dataset/test-of-time","name":"Test-of-Time","full_name":"Test of Time Synthetic Video Dataset","description_markdown":"The goal of this dataset is to probe video-language models for understanding of simple temporal relations like \"before\" and \"after\". The dataset  is only meant to be an evaluation set and not a training set. \r\n\r\nContents:\r\n1. The dataset has synthetic videos which consists of a pair of shapes appearing gradually. For example, video for the caption \"a red circle appears after a yellow circle\" will first show a \"yellow circle\" appear and then a \"red circle\" appear. The model has to determine the right caption in comparison with a distractor caption \"a yellow circle appears after a red circle\". Note that this distractor caption has the same set of words but in a different order, motivated by the Winograd schema.\r\n2. The dataset also has a control set in which videos only have a single event, e.g., \"a red circle appears\". Note that this is a control task to ensure that these videos are not out-of-distribution for a given video model.\r\nA time-aware model shall perform perfectly well on both sets.  A space-aware model that is not time-aware shall perform poorly on the temporal task while performing perfectly on the control task.","description_withheld":null,"homepage":"https://bpiyush.github.io/testoftime-website/","introduced_date":"2023-01-05","introduced_date_note":null,"introduced_by":{"paper":"/paper/test-of-time-instilling-video-language-models","title":"Test of Time: Instilling Video-Language Models with a Sense of Time","first_author":"Piyush Bagad","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Video-Text Retrieval","url":"/task/video-text-retrieval","datasets_with_task":"/datasets/task/video-text-retrieval"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Test-of-Time"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/video-text-retrieval-on-test-of-time","task":"Video-Text Retrieval","dataset_variant":"Test-of-Time","rows":4,"metrics":["2-Class Accuracy"],"first_row_in_archive_order":{"model":"Video-LLAMA","paper":"/paper/video-llama-an-instruction-tuned-audio-visual","metrics":{"2-Class Accuracy":"88.33"},"code_links":[{"title":"damo-nlp-sg/video-llama","url":"https://github.com/damo-nlp-sg/video-llama"},{"title":"damo-nlp-sg/videollama2","url":"https://github.com/damo-nlp-sg/videollama2"},{"title":"damo-nlp-sg/videollama3","url":"https://github.com/damo-nlp-sg/videollama3"},{"title":"xinding-sys/StreamMind","url":"https://github.com/xinding-sys/StreamMind"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/timechat-a-time-sensitive-multimodal-large","title":"TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding","date":"2023-12-04","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":11,"samples_ran":7,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/videoprompter-an-ensemble-of-foundational","title":"Videoprompter: an ensemble of foundational models for zero-shot video understanding","date":"2023-10-23","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/video-llama-an-instruction-tuned-audio-visual","title":"Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding","date":"2023-06-05","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":25,"samples_ran":18,"samples_unverified":7,"pointer_only_for_licence":9,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/test-of-time-instilling-video-language-models","title":"Test of Time: Instilling Video-Language Models with a Sense of Time","date":"2023-01-05","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":1,"samples_unverified":9,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":46,"samples_ran":26,"samples_unverified":20,"pointer_only_for_licence":9,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}