{"url":"/dataset/utsd","name":"UTSD","full_name":"Unified Time Series Dataset","description_markdown":"Unified Time Series Dataset (UTSD) includes 7 domains with up to 1 billion time points with hierarchical capacities to facilitate research of large models in the field of time series. It is meticulously assembled from a blend of publicly accessible online data repositories and empirical data derived from real-world machine operations. We analyze each dataset within the collection, examining the time series through the lenses of stationarity and forecastability to allows us to characterize the level of complexity inherent to each dataset.\r\n\r\nAll datasets are classified into seven distinct domains by their source: Energy, Environment, Health, Internet of Things (IoT), Nature, Transportation, and Web with diverse sampling frequencies. UTSD is constructed with hierarchical capacities, namely UTSD-1G, UTSD-2G, UTSD-4G, and UTSD-12G, where each smaller dataset is a subset of the larger ones. A larger subset means greater data difficulty and diversity, allowing you to conduct detailed scaling experiments.\r\n\r\nSee the [paper](https://arxiv.org/pdf/2402.02368) and [codebase](https://github.com/thuml/Large-Time-Series-Model) for more information.","description_withheld":null,"homepage":"https://huggingface.co/datasets/thuml/UTSD","introduced_date":"2024-02-04","introduced_date_note":null,"introduced_by":{"paper":"/paper/timer-transformers-for-time-series-analysis","title":"Timer: Generative Pre-trained Transformers Are Large Time Series Models","first_author":"Yong liu","url":null},"license":null,"modalities":[{"name":"Time series","url":"/datasets/modality/time-series"}],"tasks":[{"name":"Anomaly Detection","url":"/task/anomaly-detection","datasets_with_task":"/datasets/task/anomaly-detection"},{"name":"Time Series Forecasting","url":"/task/time-series-forecasting","datasets_with_task":"/datasets/task/time-series-forecasting"},{"name":"Imputation","url":"/task/imputation","datasets_with_task":"/datasets/task/imputation"}],"languages":[],"variants":["UTSD"],"data_loaders":[],"num_papers_in_archive":6,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}