{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/shielding-googles-language-toxicity-model","title":"Shielding Google's language toxicity model against adversarial attacks","arxiv_id":"1801.01828","date":"2018-01-05","proceeding":null,"authors":["Nestor Rodriguez","Sergio Rojas-Galeano"],"abstract":"Lack of moderation in online communities enables participants to incur in\npersonal aggression, harassment or cyberbullying, issues that have been\naccentuated by extremist radicalisation in the contemporary post-truth politics\nscenario. This kind of hostility is usually expressed by means of toxic\nlanguage, profanity or abusive statements. Recently Google has developed a\nmachine-learning-based toxicity model in an attempt to assess the hostility of\na comment; unfortunately, it has been suggested that said model can be deceived\nby adversarial attacks that manipulate the text sequence of the comment. In\nthis paper we firstly characterise such adversarial attacks as using\nobfuscation and polarity transformations. The former deceives by corrupting\ntoxic trigger content with typographic edits, whereas the latter deceives by\ngrammatical negation of the toxic content. Then, we propose a two--stage\napproach to counter--attack these anomalies, bulding upon a recently proposed\ntext deobfuscation method and the toxicity scoring model. Lastly, we conducted\nan experiment with approximately 24000 distorted comments, showing how in this\nway it is feasible to restore toxicity of the adversarial variants, while\nincurring roughly on a twofold increase in processing time. Even though novel\nadversary challenges would keep coming up derived from the versatile nature of\nwritten language, we anticipate that techniques combining machine learning and\ntext pattern recognition methods, each one targeting different layers of\nlinguistic features, would be needed to achieve robust detection of toxic\nlanguage, thus fostering aggression--free digital interaction.","url_abs":"http://arxiv.org/abs/1801.01828v1","url_pdf":"http://arxiv.org/pdf/1801.01828v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"shielding-googles-language-toxicity-model","repo_url":"https://gitlab.com/textpatrol/gp-tp-experiment","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"negation","task_name":"Negation"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}