{"url":"/dataset/sb10k","name":"SB10k","full_name":null,"description_markdown":"The **SB10k dataset** is a valuable resource for sentiment analysis in German. Here are the key details:\r\n\r\n- **Corpus Size**: It contains approximately **10,000 German tweets**¹.\r\n- **Language**: German.\r\n- **Task**: Text classification, specifically sentiment analysis.\r\n- **Multilinguality**: Monolingual (German only).\r\n- **Size Category**: Falls within the range of **1K to 10K** examples.\r\n- **Tags**: Sentiment analysis.\r\n- **License**: **CC-BY-4.0**.\r\n\r\nThe dataset was created by annotating German tweets, with each tweet labeled by **three annotators**. Researchers have used SB10k to benchmark various machine learning classifiers, including convolutional neural networks (CNNs) and feature-based support vector machines (SVMs) for sentiment analysis²³.\r\n\r\n(1) Alienmaster/SB10k · Datasets at Hugging Face. https://huggingface.co/datasets/Alienmaster/SB10k.\r\n(2) A Twitter Corpus and Benchmark Resources for German Sentiment Analysis. https://aclanthology.org/W17-1106/.\r\n(3) A Twitter Corpus and Benchmark Resources for German Sentiment Analysis. https://aclanthology.org/W17-1106.pdf.\r\n(4) undefined. http://t.co/9rhta65MSx.\r\n(5) undefined. http://t.co/G84qcIGk7k.\r\n(6) undefined. http://t.co/LvwyZgew4Q.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Alienmaster/SB10k","introduced_date":"2017-04-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/a-twitter-corpus-and-benchmark-resources-for","title":"A Twitter Corpus and Benchmark Resources for German Sentiment Analysis","first_author":"Mark Cieliebak","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["SB10k"],"data_loaders":[],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}