{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/revisiting-unsupervised-learning-for-defect","title":"Revisiting Unsupervised Learning for Defect Prediction","arxiv_id":"1703.00132","date":"2017-03-01","proceeding":null,"authors":["Wei Fu","Tim Menzies"],"abstract":"Collecting quality data from software projects can be time-consuming and\nexpensive. Hence, some researchers explore \"unsupervised\" approaches to quality\nprediction that does not require labelled data. An alternate technique is to\nuse \"supervised\" approaches that learn models from project data labelled with,\nsay, \"defective\" or \"not-defective\". Most researchers use these supervised\nmodels since, it is argued, they can exploit more knowledge of the projects.\n  At FSE'16, Yang et al. reported startling results where unsupervised defect\npredictors outperformed supervised predictors for effort-aware just-in-time\ndefect prediction. If confirmed, these results would lead to a dramatic\nsimplification of a seemingly complex task (data mining) that is widely\nexplored in the software engineering literature.\n  This paper repeats and refutes those results as follows. (1) There is much\nvariability in the efficacy of the Yang et al. predictors so even with their\napproach, some supervised data is required to prune weaker predictors away.\n(2)Their findings were grouped across $N$ projects. When we repeat their\nanalysis on a project-by-project basis, supervised predictors are seen to work\nbetter.\n  Even though this paper rejects the specific conclusions of Yang et al., we\nstill endorse their general goal. In our our experiments, supervised predictors\ndid not perform outstandingly better than unsupervised ones for effort-aware\njust-in-time defect prediction. Hence, they may indeed be some combination of\nunsupervised learners to achieve comparable performance to supervised ones. We\ntherefore encourage others to work in this promising area.","url_abs":"http://arxiv.org/abs/1703.00132v2","url_pdf":"http://arxiv.org/pdf/1703.00132v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"revisiting-unsupervised-learning-for-defect","repo_url":"https://github.com/WeiFoo/RevisitUnsupervised","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"prediction","task_name":"Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}