{"url":"/dataset/rocstories","name":"ROCStories","full_name":null,"description_markdown":"**ROCStories** is a collection of commonsense short stories. The corpus consists of 100,000 five-sentence stories. Each story logically follows everyday topics created by Amazon Mechanical Turk workers. These stories contain a variety of commonsense causal and temporal relations between everyday events. Writers also develop an additional 3,742 Story Cloze Test stories which contain a four-sentence-long body and two candidate endings. The endings were collected by asking Mechanical Turk workers to write both a right ending and a wrong ending after eliminating original endings of given short stories. Both endings were required to make logical sense and include at least one character from the main story line. The published ROCStories dataset is constructed with ROCStories as a training set that includes 98,162 stories that exclude candidate wrong endings, an evaluation set, and a test set, which have the same structure (1 body + 2 candidate endings) and a size of 1,871.\r\n\r\nSource: [Incorporating Structured Commonsense Knowledge in Story Completion](https://arxiv.org/abs/1811.00625)\r\nImage Source: [https://cs.rochester.edu/nlp/rocstories/](https://cs.rochester.edu/nlp/rocstories/)","description_withheld":null,"homepage":"https://cs.rochester.edu/nlp/rocstories/","introduced_date":"2016-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/a-corpus-and-cloze-evaluation-for-deeper","title":"A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories","first_author":"Nasrin Mostafazadeh","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Generation","url":"/task/text-generation","datasets_with_task":"/datasets/task/text-generation"},{"name":"Emotion Classification","url":"/task/emotion-classification","datasets_with_task":"/datasets/task/emotion-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ROCStories"],"data_loaders":[{"repo":"https://github.com/tensorflow/datasets","url":"https://www.tensorflow.org/datasets/catalog/story_cloze","frameworks":["tf","jax"]}],"num_papers_in_archive":152,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-generation-on-rocstories","task":"Text Generation","dataset_variant":"ROCStories","rows":4,"metrics":["BLEU-1","Perplexity"],"first_row_in_archive_order":{"model":"Beam search + A*esque (beam)","paper":"/paper/neurologic-a-esque-decoding-constrained-text","metrics":{"BLEU-1":"34.4","Perplexity":"2.14"},"code_links":[{"title":"GXimingLu/a_star_neurologic","url":"https://github.com/GXimingLu/a_star_neurologic"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/emotion-classification-on-rocstories","task":"Emotion Classification","dataset_variant":"ROCStories","rows":2,"metrics":["F1"],"first_row_in_archive_order":{"model":"Semi-supervision","paper":"/paper/modeling-label-semantics-for-predicting","metrics":{"F1":"65.88"},"code_links":[{"title":"StonyBrookNLP/emotion-label-semantics","url":"https://github.com/StonyBrookNLP/emotion-label-semantics"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/neurologic-a-esque-decoding-constrained-text","title":"NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics","date":"2021-12-16","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/modeling-label-semantics-for-predicting","title":"Modeling Label Semantics for Predicting Emotional Reactions","date":"2020-06-09","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/modeling-naive-psychology-of-characters-in","title":"Modeling Naive Psychology of Characters in Simple Commonsense Stories","date":"2018-05-16","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}