{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/identifying-experts-in-software-libraries-and","title":"Identifying Experts in Software Libraries and Frameworks among GitHub Users","arxiv_id":"1903.08113","date":"2019-03-19","proceeding":null,"authors":["Joao Eduardo Montandon","Luciana Lourdes Silva","Marco Tulio Valente"],"abstract":"Software development increasingly depends on libraries and frameworks to\nincrease productivity and reduce time-to-market. Despite this fact, we still\nlack techniques to assess developers expertise in widely popular libraries and\nframeworks. In this paper, we evaluate the performance of unsupervised (based\non clustering) and supervised machine learning classifiers (Random Forest and\nSVM) to identify experts in three popular JavaScript libraries: facebook/react,\nmongodb/node-mongodb, and socketio/socket.io. First, we collect 13 features\nabout developers activity on GitHub projects, including commits on source code\nfiles that depend on these libraries. We also build a ground truth including\nthe expertise of 575 developers on the studied libraries, as self-reported by\nthem in a survey. Based on our findings, we document the challenges of using\nmachine learning classifiers to predict expertise in software libraries, using\nfeatures extracted from GitHub. Then, we propose a method to identify library\nexperts based on clustering feature data from GitHub; by triangulating the\nresults of this method with information available on Linkedin profiles, we show\nthat it is able to recommend dozens of GitHub users with evidences of being\nexperts in the studied JavaScript libraries. We also provide a public dataset\nwith the expertise of 575 developers on the studied libraries.","url_abs":"http://arxiv.org/abs/1903.08113v1","url_pdf":"http://arxiv.org/pdf/1903.08113v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"identifying-experts-in-software-libraries-and","repo_url":"https://github.com/castor-software/oss-graph-metrics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}