{"url":"/dataset/shaved-ice-snowflake-vm-demand-dataset","name":"Shaved Ice Snowflake VM Demand Dataset","full_name":"Snowflake Dataset for \"Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads\" paper","description_markdown":"This repository contains documentation for the dataset that accompanies our\r\n[ICPE 2025](https://icpe2025.spec.org/) paper, \"Shaved Ice: Optimal Compute Resource Commitments for\r\nDynamic Multi-Cloud Workloads\".  It also includes example [R](http://www.r-project.org) and Python notebooks to\r\nread and visualize the data, including scripts to reproduce the\r\nfigures and analysis results in the paper.\r\n\r\n[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.15015992.svg)](https://doi.org/10.5281/zenodo.15015992)\r\n\r\nThis project is archived on [Zenodo](https://zenodo.org/), an open-access repository, to ensure long-term reproducibility of the research.\r\n\r\n## Dataset\r\n\r\nThe dataset contains normalized and obfuscated hourly data about VM demand in four example Snowflake deployments over a period of 3 years from 2/1/2021 to 1/30/2024.\r\nEach hour includes (type of VM, region, number of VMs of that type) used at that time.\r\nThis dataset is available in both [compressed CSV](./hourly_normalized.csv.gz) and [Parquet](./hourly_normalized.parquet) formats.\r\n\r\n### Schema\r\n\r\n* *Timestamp*: An hourly timestamp for the record.\r\n* *VM Type*: This field is obfuscated with the precise VM identifier from the Cloud Service Provider mapped into a capital letter.\r\n* *Region*: The region where the VM was deployed.  This field is obfuscated with the precise region name from the Cloud Service Provider mapped into a number between 1 and 4.\r\n* *Count*: The number of VMs of the specified type, region, and hour.  This field is normalized such that the largest type, region, hour tuple is set to 1000 in each region and other values are scaled linearly to the nearest whole number.\r\n\r\n## Potential Use Cases\r\n\r\nProvides realistic industry dataset for further research into cloud demand forecasting, commitment optimization, and capacity planning.","description_withheld":null,"homepage":"https://github.com/Snowflake-Labs/shavedice-dataset","introduced_date":"2025-03-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/shaved-ice-optimal-compute-resource","title":"Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads","first_author":null,"url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Graphs","url":"/datasets/modality/graphs"},{"name":"Time series","url":"/datasets/modality/time-series"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Shaved Ice Snowflake VM Demand Dataset"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}