{"url":"/dataset/capriccio","name":"Capriccio","full_name":"Sentiment Analysis + Data Drift","description_markdown":"Capriccio is a sentiment classification dataset on tweets that simulates data drift.\r\nIt is created by slicing the Sentiment140 dataset ([homepage](http://help.sentiment140.com/home), [Huggingface datasets](https://huggingface.co/datasets/sentiment140)) with a sliding window of 500,000 tweets, resulting in 38 slices.\r\nThus, each slice can be used to represent the training/validation dataset of a sentiment classification model that is re-trained every day.\r\nEach slice has 425,000 tweets for training (file named `%d_train.json`) and 75,000 tweets for validation (file named `%d_val.json`).\r\n\r\nThe name comes from the adjective *capricious*.","description_withheld":null,"homepage":"https://github.com/SymbioticLab/Zeus/tree/master/capriccio","introduced_date":"2022-08-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/zeus-understanding-and-optimizing-gpu-energy","title":"Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training","first_author":"Jie You","url":null},"license":{"name":"Apache-2.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Sentiment Analysis","url":"/task/sentiment-analysis","datasets_with_task":"/datasets/task/sentiment-analysis"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Capriccio"],"data_loaders":[{"repo":"https://github.com/symbioticlab/zeus","url":"https://github.com/SymbioticLab/Zeus/tree/master/capriccio","frameworks":["pytorch"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}