{"url":"/dataset/korean-multi-label-hate-speech-dataset","name":"K-MHaS: Korean Multi-label Hate Speech Dataset","full_name":null,"description_markdown":"Korean Multi-label Hate Speech Dataset\r\n\r\nWe introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns.\r\n\r\n* consisting of 109,692 utterances from Korean online news comments, labeled with 8 fine-grained hate speech classes.\r\n* data collection period: between January 2018 and June 2020.\r\n\r\n* providing (a) binary classification and (b) multi-label classification from 1(one) to 4(four) labels.\r\n* (a) binary classification: Hate Speech or Not Hate Speech\r\n* (b) fine-grained classification: Politics, Origin, Physical, Age, Gender, Religion, Race, and Profanity.\r\n\r\nFor the fine-grained classification, a Hate Speech class from the binary classification is broken down into eight classes, associated with the hate speech category.","description_withheld":null,"homepage":"https://github.com/adlnlp/K-MHaS","introduced_date":"2022-08-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/k-mhas-a-multi-label-hate-speech-detection","title":"K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment","first_author":"Jean Lee","url":null},"license":{"name":"cc-by-sa-4.0","url":"https://creativecommons.org/licenses/by-sa/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Hate Speech Detection","url":"/task/hate-speech-detection","datasets_with_task":"/datasets/task/hate-speech-detection"},{"name":"Toxic Comment Classification","url":"/task/toxic-comment-classification","datasets_with_task":"/datasets/task/toxic-comment-classification"}],"languages":[{"name":"Korean","url":"/datasets/language/korean"}],"variants":["K-MHaS: Korean Multi-label Hate Speech Dataset"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/trueorfalse441/korean_hate_speech_copy","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/jeanlee/kmhas_korean_hate_speech","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/adlnlp/K-MHaS","url":"https://github.com/adlnlp/K-MHaS","frameworks":["tf","pytorch"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}