{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/detecting-sexism-in-german-online-newspaper","title":"Detecting Sexism in German Online Newspaper Comments with Open-Source Text Embeddings (Team GDA, GermEval2024 Shared Task 1: GerMS-Detect, Subtasks 1 and 2, Closed Track)","arxiv_id":"2409.10341","date":"2024-09-16","proceeding":null,"authors":["Florian Bremm","Patrick Gustav Blaneck","Tobias Bornheim","Niklas Grieger","Stephan Bialonski"],"abstract":"Sexism in online media comments is a pervasive challenge that often manifests subtly, complicating moderation efforts as interpretations of what constitutes sexism can vary among individuals. We study monolingual and multilingual open-source text embeddings to reliably detect sexism and misogyny in German-language online comments from an Austrian newspaper. We observed classifiers trained on text embeddings to mimic closely the individual judgements of human annotators. Our method showed robust performance in the GermEval 2024 GerMS-Detect Subtask 1 challenge, achieving an average macro F1 score of 0.597 (4th place, as reported on Codabench). It also accurately predicted the distribution of human annotations in GerMS-Detect Subtask 2, with an average Jensen-Shannon distance of 0.301 (2nd place). The computational efficiency of our approach suggests potential for scalable applications across various languages and linguistic contexts.","url_abs":"https://arxiv.org/abs/2409.10341v2","url_pdf":"https://arxiv.org/pdf/2409.10341v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"detecting-sexism-in-german-online-newspaper","repo_url":"https://github.com/dslaborg/germeval2024","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"germeval2024-shared-task-1-subtask-1","task_name":"GermEval2024 Shared Task 1 Subtask 1"},{"task_slug":"germeval2024-shared-task-1-subtask-2","task_name":"GermEval2024 Shared Task 1 Subtask 2"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/germeval2024-shared-task-1-subtask-1-on-germs","task":"GermEval2024 Shared Task 1 Subtask 1","dataset":"GerMS-AT","model":"mE5-large-SVM","rank_in_archive_order":1,"of":1,"metrics":{"Macro F1":"0.597"},"uses_additional_data":false},{"leaderboard":"/sota/germeval2024-shared-task-1-subtask-2-on-germs","task":"GermEval2024 Shared Task 1 Subtask 2","dataset":"GerMS-AT","model":"GBERT-large-SVM","rank_in_archive_order":1,"of":1,"metrics":{"Jensen-Shannon distance":"0.301"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}