{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-dictionary-based-approach-to-racism","title":"A Dictionary-based Approach to Racism Detection in Dutch Social Media","arxiv_id":"1608.08738","date":"2016-08-31","proceeding":null,"authors":["Stéphan Tulkens","Lisa Hilte","Elise Lodewyckx","Ben Verhoeven","Walter Daelemans"],"abstract":"We present a dictionary-based approach to racism detection in Dutch social\nmedia comments, which were retrieved from two public Belgian social media sites\nlikely to attract racist reactions. These comments were labeled as racist or\nnon-racist by multiple annotators. For our approach, three discourse\ndictionaries were created: first, we created a dictionary by retrieving\npossibly racist and more neutral terms from the training data, and then\naugmenting these with more general words to remove some bias. A second\ndictionary was created through automatic expansion using a \\texttt{word2vec}\nmodel trained on a large corpus of general Dutch text. Finally, a third\ndictionary was created by manually filtering out incorrect expansions. We\ntrained multiple Support Vector Machines, using the distribution of words over\nthe different categories in the dictionaries as features. The best-performing\nmodel used the manually cleaned dictionary and obtained an F-score of 0.46 for\nthe racist class on a test set consisting of unseen Dutch comments, retrieved\nfrom the same sites used for the training set. The automated expansion of the\ndictionary only slightly boosted the model's performance, and this increase in\nperformance was not statistically significant. The fact that the coverage of\nthe expanded dictionaries did increase indicates that the words that were\nautomatically added did occur in the corpus, but were not able to meaningfully\nimpact performance. The dictionaries, code, and the procedure for requesting\nthe corpus are available at: https://github.com/clips/hades","url_abs":"http://arxiv.org/abs/1608.08738v1","url_pdf":"http://arxiv.org/pdf/1608.08738v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-dictionary-based-approach-to-racism","repo_url":"https://github.com/clips/hades","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1608.08738","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}