{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hate-speech-detection-a-solved-problem-the","title":"Hate Speech Detection: A Solved Problem? The Challenging Case of Long Tail on Twitter","arxiv_id":"1803.03662","date":"2018-02-27","proceeding":null,"authors":["Ziqi Zhang","Lei Luo"],"abstract":"In recent years, the increasing propagation of hate speech on social media\nand the urgent need for effective counter-measures have drawn significant\ninvestment from governments, companies, and researchers. A large number of\nmethods have been developed for automated hate speech detection online. This\naims to classify textual content into non-hate or hate speech, in which case\nthe method may also identify the targeting characteristics (i.e., types of\nhate, such as race, and religion) in the hate speech. However, we notice\nsignificant difference between the performance of the two (i.e., non-hate v.s.\nhate). In this work, we argue for a focus on the latter problem for practical\nreasons. We show that it is a much more challenging task, as our analysis of\nthe language in the typical datasets shows that hate speech lacks unique,\ndiscriminative features and therefore is found in the 'long tail' in a dataset\nthat is difficult to discover. We then propose Deep Neural Network structures\nserving as feature extractors that are particularly effective for capturing the\nsemantics of hate speech. Our methods are evaluated on the largest collection\nof hate speech datasets based on Twitter, and are shown to be able to\noutperform the best performing method by up to 5 percentage points in\nmacro-average F1, or 8 percentage points in the more challenging case of\nidentifying hateful content.","url_abs":"http://arxiv.org/abs/1803.03662v2","url_pdf":"http://arxiv.org/pdf/1803.03662v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hate-speech-detection-a-solved-problem-the","repo_url":"https://github.com/ziqizhang/chase","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"hate-speech-detection","task_name":"Hate Speech Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.03662","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1803.03662"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ziqizhang/chase","reach":null}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"43e9ba10dd4b43ae","entry":"build_pretrained_embedding_matrix","repo":"ziqizhang/chase","repo_kind":"official","path":"python/src/ml/classifier_dnn.py","file_url":"https://github.com/ziqizhang/chase/blob/HEAD/python/src/ml/classifier_dnn.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"43e9ba10dd4b43ae"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}