Home › Datasets › task › Hate Speech Detection
Hate Speech Detection datasets
archive 2025-07-28
47 datasets carry the task tag "Hate Speech Detection" (the task itself: Hate Speech Detection), ordered by the archive's paper count. Page 1 of 1: 47 shown of 47. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Hate Speech Detection datasets 1–47 of 47
Dataset of hate speech annotated on Internet forum posts in English at sentence-level.
180 papers · 1 benchmark
At the end of 2017 the Civil Comments platform shut down and chose make their ~2m public comments from their platform available in a lasting open archive so that researchers could understand and improve civility in online conversations for…
156 papers · 1 benchmark
OLID (Offensive Language Identification Dataset)
The OLID is a hierarchical dataset to identify the type and the target of offensive texts in social media.
152 papers · 1 benchmark
Covers multiple aspects of the issue.
105 papers · 2 benchmarks
A large-scale and machine-generated dataset of 274,186 toxic and benign statements about 13 minority groups.
85 papers · 0 benchmarks
HatEval (SemEval 2019 Task 5: Multilingual Detection of Hate Speech Against Immigrants and Women in Twitter)
Hate Speech is commonly defined as any communication that disparages a person or a group on the basis of some characteristic such as race, color, ethnicity, gender, sexual orientation, nationality, religion, or other characteristics.
67 papers · 1 benchmark
HSOL is a dataset for hate speech detection.
67 papers · 0 benchmarks
The Implicit Hate corpus is a dataset for hate speech detection with fine-grained labels for each message and its implication.
58 papers · 0 benchmarks
ETHOS (multi-labEl haTe speecH detectiOn dataSet)
ETHOS is a hate speech detection dataset.
20 papers · 2 benchmarks
Existing hate speech datasets contain only textual data.
15 papers · 0 benchmarks
ViHSD (Vietnamese Hate Speech Detection Dataset)
This dataset contains 33,400 annotated comments used for hate speech detection on social network sites.
12 papers · 0 benchmarks
Introduces three datasets of expressing hate, commonly used topics, and opinions for hate speech detection, document classification, and sentiment analysis, respectively.
6 papers · 0 benchmarks
SWSR (Sina Weibo Sexism Review)
The Sina Weibo Sexism Review (SWSR) dataset is a dataset to research online sexism in Chinese.
6 papers · 0 benchmarks
A corpus of Offensive Language and Hate Speech Detection for Danish This DKhate dataset contains 3600 comments from the web annotated for offensive language, following the Zampieri et al.
5 papers · 1 benchmark
DeToxy (DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances)
DeToxy is a publicly available toxicity annotated dataset for the English language.
5 papers · 0 benchmarks
ToLD-Br (Toxic Language Detection for Brazilian Portuguese)
The Toxic Language Detection for Brazilian Portuguese (ToLD-Br) is a dataset with tweets in Brazilian Portuguese annotated according to different toxic aspects.
5 papers · 1 benchmark
Hate speech has become one of the most significant issues in modern society, with implications in both the online and offline worlds.
4 papers · 1 benchmark
A new multilingual multi-aspect hate speech analysis dataset and use it to test the current state-of-the-art multilingual multitask learning approaches.
4 papers · 0 benchmarks
HatemojiCheck is a test suite for detecting emoji-based hate of 3,930 test cases covering seven functionalities of emoji-based hate and six identities.
3 papers · 0 benchmarks
Presents 9.4K manually labeled entertainment news comments for identifying Korean toxic speech, collected from a widely used online news platform in Korea.
3 papers · 0 benchmarks
AraCOVID19-MFH (AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset)
AraCOVID19-MFH is a manually annotated multi-label Arabic COVID-19 fake news and hate speech detection dataset.
2 papers · 0 benchmarks
FairPrism is a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harms relating to gender and sexuality.
2 papers · 0 benchmarks
HERDPhobia is an annotated hate speech detection dataset on Fulani herders in Nigeria -- in three languages: English, Nigerian-Pidgin, and Hausa.
2 papers · 0 benchmarks
LatamXIX (19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction)
A novel dataset of 19th-century Latin American press texts, which addresses the lack of specialized corpora for historical and linguistic analysis in this region.
2 papers · 0 benchmarks
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
chinahate dataset contains a total of 2,172,333 tweets hashtagged #china posted during the time it was collected.
1 paper · 0 benchmarks
The dataset contains 7,601 Gab posts classified on three different aspects: abuse presence or not, abuse severity and abuse target.
1 paper · 0 benchmarks
CoRAL dataset (CoRAL: a Context-aware Croatian Abusive Language Dataset)
CoRAL is a language and culturally aware Croatian Abusive dataset covering phenomena of implicitness and reliance on local and global context.
1 paper · 0 benchmarks
A large dataset of over 18,000,000 English tweets posted by ∼7K echo users was constructed in the following manner: 1.
1 paper · 0 benchmarks
Echo Corpus (Arviv et al, 2021) infused with information from KnowledJe (Halevy, 2023).
1 paper · 0 benchmarks
HS-BAN is a binary class hate speech (HS) dataset in Bangla language consisting of more than 50,000 labeled comments, including 40.17% hate and rest are non hate speech.
1 paper · 0 benchmarks
Multi-Modal Hate Speech Detection with Graph Context.
1 paper · 0 benchmarks
Dataset Description - Paper: TBC - Point of Contact: Josh McGiff (Josh.McGiff@ul.ie) Dataset Summary This dataset was developed to address the significant gap in online hate speech detection, particularly focusing on homophobia, which is…
1 paper · 0 benchmarks
Korean Multi-label Hate Speech Dataset We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns.
1 paper · 0 benchmarks
We introduce KnowledJe, an English-language knowledge graph of antisemitic history and language from the 20th century to the present.
1 paper · 0 benchmarks
APEACH is the first crowd-generated Korean evaluation dataset for hate speech detection.
1 paper · 0 benchmarks
L3Cube-MahaCorpus is a Marathi monolingual data set scraped from different internet sources.
1 paper · 0 benchmarks
M-Phasis (A Feature-Based Corpus of Hate Online)
A corpus of 9k German and French user comments collected from migration-related news articles.
1 paper · 0 benchmarks
NJH is a dataset of over 40,000 tweets about immigration from the US and UK, annotated with six labels for different aspects of incivility and intolerance.
1 paper · 0 benchmarks
Peer to Peer Hate is a comprehensive hate speech dataset capturing various types of hate.
1 paper · 0 benchmarks
SHAJ (Spoken Hate in the Albanian Jargon)
This is an abusive/offensive language detection dataset for Albanian.
1 paper · 1 benchmark
With the rise of social media, user-generated content has surged, and hate speech has proliferated.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The ComMA Dataset v0.2 is a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
1 paper · 0 benchmarks
ViTHSD (Vietnamese Targeted-Hate-Speech-Detection)
A Vietnamese dataset for hate speech detection by the specific target.
1 paper · 0 benchmarks
This is a high-quality dataset of annotated posts sampled from social media posts and annotated for misogyny.
1 paper · 1 benchmark
Arabic multi-dialectal hate speech dataset.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.