{"url":"/dataset/sentimentarcs-sentiment-reference-corpus-for","name":"SentimentArcs: Sentiment Reference Corpus for Novels","full_name":"SentimentArcs: Sentiment Reference Corpus for Novels","description_markdown":"SentimentArcs’ reference corpus for novels consists of 25 narratives selected to create a diverse set of well recognized novels that can serve as a benchmark for future studies. The composition of the corpora was limited by the effect of copyright laws as well as historical imbalances. Most works were obtained from US and Australian Gutenberg Projects. The corpora is expected to grow in size and diversity over time.  \r\n\r\nSeveral dimensions of diversity were considered for inclusion including popularity, period, genre, topic, style and author diversity. The first version of our corpus includes only English, although Proust and Homer are included in translation. SentimentArcs has processed a larger set of novels, including some in foreign languages. The initial reference corpus is in English since performance across all ensemble models was uneven in less resourced languages\r\n\r\nIn sum, the corpora includes (1) the two most popular novels on Gutenberg.org (Project Gutenberg, 2021b), (2) eight of the fifteen most assigned novels at top US universities (EAB, 2021), and (3) three works that have sold over 20 million copies (Books, 2021). There are eight works by women, two by African-Americans and five works by two LGBTQ authors. Britain leads with 15 authors followed by 6 Americans and one each from France, Russia, North Africa and Ancient Greece.","description_withheld":null,"homepage":"https://github.com/jon-chun/sentimentarcs_notebooks/","introduced_date":"2021-10-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/sentimentarcs-a-novel-method-for-self","title":"SentimentArcs: A Novel Method for Self-Supervised Sentiment Analysis of Time Series Shows SOTA Transformers Can Struggle Finding Narrative Arcs","first_author":"Jon Chun","url":null},"license":{"name":"MIT","url":"https://github.com/jon-chun/sentimentarcs_notebooks/blob/main/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Sentiment Analysis","url":"/task/sentiment-analysis","datasets_with_task":"/datasets/task/sentiment-analysis"},{"name":"Twitter Sentiment Analysis","url":"/task/twitter-sentiment-analysis","datasets_with_task":"/datasets/task/twitter-sentiment-analysis"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SentimentArcs: Sentiment Reference Corpus for Novels"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}