{"url":"/dataset/hatespeechcorpus","name":"HateSpeechCorpus","full_name":null,"description_markdown":"HateSpeechCorpus is a dataset collected from Twitter, consisting of 3003 tweets. It has been meticulously annotated by three speech-language pathology graduate students, ensuring high-quality labeling of hate speech. This dataset is invaluable for researchers and practitioners working on hate speech detection and natural language processing.\r\n\r\n**Source:** Twitter  \r\n**Length:** 3003 tweets  \r\n**Annotators:** Three speech-language pathology graduate students  \r\n\r\n## Paper\r\nFor a detailed investigation of annotator bias in LLMs for hate speech detection, refer to the associated paper: [Investigating Annotator Bias in Large Language Models for Hate Speech Detection](https://arxiv.org/abs/2406.11109).\r\n\r\n## Citation\r\nIf you use this dataset in your research, please cite the following paper:\r\n\r\n@misc{das2024investigating,\r\ntitle={Investigating Annotator Bias in Large Language Models for Hate Speech Detection},\r\nauthor={Amit Das and Zheng Zhang and Fatemeh Jamshidi and Vinija Jain and Aman Chadha and Nilanjana Raychawdhary and Mary Sandage and Lauramarie Pope and Gerry Dozier and Cheryl Seals},\r\nyear={2024},\r\neprint={2406.11109},\r\narchivePrefix={arXiv},\r\nprimaryClass={cs.CL}\r\n}\r\n\r\n# Load the dataset\r\ndf = pd.read_csv('HateSpeechCorpus.csv')\r\n\r\n# Display the first few rows\r\nprint(df.head())\r\n\r\n## License: CC BY\r\n\r\nContact Information: azd0123@auburn.edu\r\n---","description_withheld":null,"homepage":"","introduced_date":"2024-06-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/investigating-annotator-bias-in-large","title":"Investigating Annotator Bias in Large Language Models for Hate Speech Detection","first_author":"Amit Das","url":null},"license":{"name":"CC BY","url":null},"modalities":[],"tasks":[],"languages":[],"variants":["HateSpeechCorpus"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}