{"url":"/dataset/tupye-dataset","name":"TuPyE-Dataset","full_name":"Portuguese Hate Speech Expanded Dataset","description_markdown":"TuPyE, an enhanced iteration of TuPy, encompasses a compilation of 43,668 meticulously annotated documents specifically selected for the purpose of hate speech detection within diverse social network contexts. This augmented dataset integrates supplementary annotations and amalgamates with datasets sourced from Fortuna et al. (2019), Leite et al. (2020), and Vargas et al. (2022), complemented by an infusion of 10,000 original documents from the TuPy-Dataset.\r\n\r\nIn light of the constrained availability of annotated data in Portuguese pertaining to the English language, TuPyE is committed to the expansion and enhancement of existing datasets. This augmentation serves to facilitate the development of advanced hate speech detection models through the utilization of machine learning (ML) and natural language processing (NLP) techniques.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Silly-Machine/TuPyE-Dataset","introduced_date":"2023-12-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/tupy-e-detecting-hate-speech-in-brazilian","title":"TuPy-E: detecting hate speech in Brazilian Portuguese social media with a novel dataset and comprehensive analysis of models","first_author":"Felipe Oliveira","url":null},"license":{"name":"cc-by-4.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Binary Classification","url":"/task/binary-classification","datasets_with_task":"/datasets/task/binary-classification"},{"name":"Multilabel Text Classification","url":"/task/multilabel-text-classification","datasets_with_task":"/datasets/task/multilabel-text-classification"},{"name":"Hate Span Identification","url":"/task/hate-span-identification","datasets_with_task":"/datasets/task/hate-span-identification"}],"languages":[{"name":"Portuguese","url":"/datasets/language/portuguese"}],"variants":["TuPyE-Dataset"],"data_loaders":[{"repo":"https://github.com/Silly-Machine/TuPyE-Dataset","url":"https://huggingface.co/datasets/Silly-Machine/TuPyE-Dataset","frameworks":["tf","pytorch"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}