{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/convolutional-neural-networks-for-toxic","title":"Convolutional Neural Networks for Toxic Comment Classification","arxiv_id":"1802.09957","date":"2018-02-27","proceeding":null,"authors":["Spiros V. Georgakopoulos","Sotiris K. Tasoulis","Aristidis G. Vrahatis","Vassilis P. Plagianakos"],"abstract":"Flood of information is produced in a daily basis through the global Internet\nusage arising from the on-line interactive communications among users. While\nthis situation contributes significantly to the quality of human life,\nunfortunately it involves enormous dangers, since on-line texts with high\ntoxicity can cause personal attacks, on-line harassment and bullying behaviors.\nThis has triggered both industrial and research community in the last few years\nwhile there are several tries to identify an efficient model for on-line toxic\ncomment prediction. However, these steps are still in their infancy and new\napproaches and frameworks are required. On parallel, the data explosion that\nappears constantly, makes the construction of new machine learning\ncomputational tools for managing this information, an imperative need.\nThankfully advances in hardware, cloud computing and big data management allow\nthe development of Deep Learning approaches appearing very promising\nperformance so far. For text classification in particular the use of\nConvolutional Neural Networks (CNN) have recently been proposed approaching\ntext analytics in a modern manner emphasizing in the structure of words in a\ndocument. In this work, we employ this approach to discover toxic comments in a\nlarge pool of documents provided by a current Kaggle's competition regarding\nWikipedia's talk page edits. To justify this decision we choose to compare CNNs\nagainst the traditional bag-of-words approach for text analysis combined with a\nselection of algorithms proven to be very effective in text classification. The\nreported results provide enough evidence that CNN enhance toxic comment\nclassification reinforcing research interest towards this direction.","url_abs":"http://arxiv.org/abs/1802.09957v1","url_pdf":"http://arxiv.org/pdf/1802.09957v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"convolutional-neural-networks-for-toxic","repo_url":"https://github.com/xinzhel/kaggle-toxicity-2021","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"cloud-computing","task_name":"Cloud Computing"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"management","task_name":"Management"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"toxic-comment-classification","task_name":"Toxic Comment Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.09957","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}