{"url":"/dataset/hatebr","name":"HateBR","full_name":null,"description_markdown":"The **HateBR dataset** is a significant resource for studying offensive language and hate speech detection in Brazilian Portuguese. Here are the key details about this dataset:\r\n\r\n1. **Collection and Annotation**:\r\n   - The HateBR dataset was **collected from Brazilian Instagram comments** related to politicians.\r\n   - It was **manually annotated by specialists** who carefully labeled each comment.\r\n   - The dataset consists of **7,000 documents**.\r\n\r\n2. **Annotation Layers**:\r\n   - The HateBR dataset includes annotations at three different levels:\r\n     - **Binary Classification**: Comments are labeled as either **offensive** or **non-offensive**.\r\n     - **Offensiveness Levels**: Comments are categorized as **highly**, **moderately**, or **slightly offensive**.\r\n     - **Hate Speech Targets**: Comments are further classified into **nine** specific hate speech categories:\r\n       - Xenophobia\r\n       - Racism\r\n       - Homophobia\r\n       - Sexism\r\n       - Religious intolerance\r\n       - Partyism\r\n       - Apology for the dictatorship\r\n       - Antisemitism\r\n       - Fatphobia\r\n\r\n3. **Inter-Annotator Agreement**:\r\n   - Each comment was annotated by **three different annotators** to ensure reliability.\r\n   - The dataset achieved **high inter-annotator agreement**.\r\n\r\n4. **Baseline Performance**:\r\n   - Baseline experiments using machine learning models achieved an **F1-score of 85%**, outperforming existing baselines for Portuguese language hate speech datasets.\r\n\r\n5. **Corpus and Models**:\r\n   - The HateBR dataset includes a **corpus** of annotated comments.\r\n   - The repository contains the **best models** presented in the associated research paper.\r\n\r\n6. **File Format**:\r\n   - The `HateBr.csv` file provides four columns:\r\n     - 1st column: Instagram comments.\r\n     - 2nd column: Offensive language classification (offensive vs. non-offensive).\r\n     - 3rd column: Offensiveness level (highly, moderately, slightly offensive).\r\n     - 4th column: Hate speech classification (nine different targets).\r\n\r\nSource: Conversation with Bing, 3/16/2024\r\n(1) HateBR - Offensive Language and Hate Speech Dataset in ... - GitHub. https://github.com/franciellevargas/HateBR.\r\n(2) ruanchaves/hatebr · Datasets at Hugging Face. https://huggingface.co/datasets/ruanchaves/hatebr.\r\n(3) Papers with Code - HateBR: Large expert annotated corpus of Brazilian .... https://paperswithcode.com/paper/hatebr-large-expert-annotated-corpus-of.","description_withheld":null,"homepage":"https://github.com/franciellevargas/HateBR","introduced_date":"2021-03-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/annotating-hate-and-offenses-on-social-media","title":"HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection","first_author":"Francielle Alves Vargas","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["HateBR"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}