{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-impact-of-ir-based-classifier","title":"The Impact of IR-based Classifier Configuration on the Performance and the Effort of Method-Level Bug Localization","arxiv_id":"1806.07727","date":"2018-06-20","proceeding":null,"authors":["Tantithamthavorn Chakkrit","Abebe Surafel Lemma","Hassan Ahmed E.","Ihara Akinori","Matsumoto Kenichi"],"abstract":"Context: IR-based bug localization is a classifier that assists developers in\nlocating buggy source code entities (e.g., files and methods) based on the\ncontent of a bug report. Such IR-based classifiers have various parameters that\ncan be configured differently (e.g., the choice of entity representation).\nObjective: In this paper, we investigate the impact of the choice of the\nIR-based classifier configuration on the top-k performance and the required\neffort to examine source code entities before locating a bug at the method\nlevel. Method: We execute a large space of classifier configuration, 3,172 in\ntotal, on 5,266 bug reports of two software systems, i.e., Eclipse and Mozilla.\nResults: We find that (1) the choice of classifier configuration impacts the\ntop-k performance from 0.44% to 36% and the required effort from 4,395 to\n50,000 LOC; (2) classifier configurations with similar top-k performance might\nrequire different efforts; (3) VSM achieves both the best top-k performance and\nthe least required effort for method-level bug localization; (4) the likelihood\nof randomly picking a configuration that performs within 20% of the best top-k\nclassifier configuration is on average 5.4% and that of the least effort is on\naverage 1%; (5) configurations related to the entity representation of the\nanalyzed data have the most impact on both the top-k performance and the\nrequired effort; and (6) the most efficient classifier configuration obtained\nat the method-level can also be used at the file-level (and vice versa).\nConclusion: Our results lead us to conclude that configuration has a large\nimpact on both the top-k performance and the required effort for method-level\nbug localization, suggesting that the IR-based configuration settings should be\ncarefully selected and the required effort metric should be included in future\nbug localization studies.","url_abs":"http://arxiv.org/abs/1806.07727v1","url_pdf":"http://arxiv.org/pdf/1806.07727v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-impact-of-ir-based-classifier","repo_url":"https://github.com/SAILResearch/replication-ist_bug_localization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}