{"url":"/dataset/tot","name":"ToT","full_name":"Test of Time","description_markdown":"ToT is a benchmark for evaluating LLMs on temporal reasoning. \r\n\r\nToT is a dataset designed to assess the temporal reasoning capabilities of AI models. It comprises two key sections:\r\n\r\nToT-semantic: Measuring the semantics and logic of time understanding.\r\nToT-arithmetic: Measuring the ability to carry out time arithmetic operations.\r\n\r\nData Format\r\nThe ToT-semantic and ToT-semantic-large datasets contain the following fields:\r\n\r\nquestion: Contains the text of the question.\r\ngraph_gen_algorithm: Contains the name of the graph generator algorithm used to generate the graph.\r\nquestion_type: Corresponds to one of the 7 question types in the dataset.\r\nsorting_type: Correspons to the sorting type applied on the facts to order them.\r\nprompt: Contains the full prompt text used to evaluate LLMs on the task.\r\nlabel: Contains the ground truth answer to the question.\r\nThe ToT-arithmetic dataset contains the following fields:\r\n\r\nquestion: Contains the text of the question.\r\nquestion_type: Corresponds to one of the 7 question types in the dataset.\r\nlabel: Contains the ground truth answer to the question.","description_withheld":null,"homepage":"https://huggingface.co/datasets/baharef/ToT","introduced_date":"2024-06-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/test-of-time-a-benchmark-for-evaluating-llms","title":"Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning","first_author":"Bahare Fatemi","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["ToT"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}