{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/topic-modeling-of-public-repositories-at","title":"Topic modeling of public repositories at scale using names in source code","arxiv_id":"1704.00135","date":"2017-04-01","proceeding":null,"authors":["Vadim Markovtsev","Eiso Kant"],"abstract":"Programming languages themselves have a limited number of reserved keywords\nand character based tokens that define the language specification. However,\nprogrammers have a rich use of natural language within their code through\ncomments, text literals and naming entities. The programmer defined names that\ncan be found in source code are a rich source of information to build a high\nlevel understanding of the project. The goal of this paper is to apply topic\nmodeling to names used in over 13.6 million repositories and perceive the\ninferred topics. One of the problems in such a study is the occurrence of\nduplicate repositories not officially marked as forks (obscure forks). We show\nhow to address it using the same identifiers which are extracted for topic\nmodeling.\n  We open with a discussion on naming in source code, we then elaborate on our\napproach to remove exact duplicate and fuzzy duplicate repositories using\nLocality Sensitive Hashing on the bag-of-words model and then discuss our work\non topic modeling; and finally present the results from our data analysis\ntogether with open-access to the source code, tools and datasets.","url_abs":"http://arxiv.org/abs/1704.00135v2","url_pdf":"http://arxiv.org/pdf/1704.00135v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"topic-modeling-of-public-repositories-at","repo_url":"https://github.com/JetBrains-Research/buckwheat","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"topic-modeling-of-public-repositories-at","repo_url":"https://github.com/JetBrains-Research/sosed","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}