{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/measuring-the-reliability-of-hate-speech","title":"Measuring the Reliability of Hate Speech Annotations: The Case of the European Refugee Crisis","arxiv_id":"1701.08118","date":"2017-01-27","proceeding":null,"authors":["Björn Ross","Michael Rist","Guillermo Carbonell","Benjamin Cabrera","Nils Kurowsky","Michael Wojatzki"],"abstract":"Some users of social media are spreading racist, sexist, and otherwise\nhateful content. For the purpose of training a hate speech detection system,\nthe reliability of the annotations is crucial, but there is no universally\nagreed-upon definition. We collected potentially hateful messages and asked two\ngroups of internet users to determine whether they were hate speech or not,\nwhether they should be banned or not and to rate their degree of offensiveness.\nOne of the groups was shown a definition prior to completing the survey. We\naimed to assess whether hate speech can be annotated reliably, and the extent\nto which existing definitions are in accordance with subjective ratings. Our\nresults indicate that showing users a definition caused them to partially align\ntheir own opinion with the definition but did not improve reliability, which\nwas very low overall. We conclude that the presence of hate speech should\nperhaps not be considered a binary yes-or-no decision, and raters need more\ndetailed instructions for the annotation.","url_abs":"http://arxiv.org/abs/1701.08118v1","url_pdf":"http://arxiv.org/pdf/1701.08118v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"measuring-the-reliability-of-hate-speech","repo_url":"https://github.com/UCSM-DUE/IWG_hatespeech_public","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"hate-speech-detection","task_name":"Hate Speech Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.08118","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}