{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-case-for-being-average-a-mediocrity","title":"The Case for Being Average: A Mediocrity Approach to Style Masking and Author Obfuscation","arxiv_id":"1707.03736","date":"2017-07-12","proceeding":null,"authors":["Georgi Karadjov","Tsvetomila Mihaylova","Yasen Kiprov","Georgi Georgiev","Ivan Koychev","Preslav Nakov"],"abstract":"Users posting online expect to remain anonymous unless they have logged in,\nwhich is often needed for them to be able to discuss freely on various topics.\nPreserving the anonymity of a text's writer can be also important in some other\ncontexts, e.g., in the case of witness protection or anonymity programs.\nHowever, each person has his/her own style of writing, which can be analyzed\nusing stylometry, and as a result, the true identity of the author of a piece\nof text can be revealed even if s/he has tried to hide it. Thus, it could be\nhelpful to design automatic tools that can help a person obfuscate his/her\nidentity when writing text. In particular, here we propose an approach that\nchanges the text, so that it is pushed towards average values for some general\nstylometric characteristics, thus making the use of these characteristics less\ndiscriminative. The approach consists of three main steps: first, we calculate\nthe values for some popular stylometric metrics that can indicate authorship;\nthen we apply various transformations to the text, so that these metrics are\nadjusted towards the average level, while preserving the semantics and the\nsoundness of the text; and finally, we add random noise. This approach turned\nout to be very efficient, and yielded the best performance on the Author\nObfuscation task at the PAN-2016 competition.","url_abs":"http://arxiv.org/abs/1707.03736v2","url_pdf":"http://arxiv.org/pdf/1707.03736v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-case-for-being-average-a-mediocrity","repo_url":"https://bitbucket.org/pan2016authorobfuscation/authorobfuscation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"the-case-for-being-average-a-mediocrity","repo_url":"https://github.com/asad1996172/Obfuscation-Detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.03736","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}