{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-meta-analysis-of-the-anomaly-detection","title":"A Meta-Analysis of the Anomaly Detection Problem","arxiv_id":"1503.01158","date":"2015-03-03","proceeding":null,"authors":["Andrew Emmott","Shubhomoy Das","Thomas Dietterich","Alan Fern","Weng-Keen Wong"],"abstract":"This article provides a thorough meta-analysis of the anomaly detection\nproblem. To accomplish this we first identify approaches to benchmarking\nanomaly detection algorithms across the literature and produce a large corpus\nof anomaly detection benchmarks that vary in their construction across several\ndimensions we deem important to real-world applications: (a) point difficulty,\n(b) relative frequency of anomalies, (c) clusteredness of anomalies, and (d)\nrelevance of features. We apply a representative set of anomaly detection\nalgorithms to this corpus, yielding a very large collection of experimental\nresults. We analyze these results to understand many phenomena observed in\nprevious work. First we observe the effects of experimental design on\nexperimental results. Second, results are evaluated with two metrics, ROC Area\nUnder the Curve and Average Precision. We employ statistical hypothesis testing\nto demonstrate the value (or lack thereof) of our benchmarks. We then offer\nseveral approaches to summarizing our experimental results, drawing several\nconclusions about the impact of our methodology as well as the strengths and\nweaknesses of some algorithms. Last, we compare results against a trivial\nsolution as an alternate means of normalizing the reported performance of\nalgorithms. The intended contributions of this article are many; in addition to\nproviding a large publicly-available corpus of anomaly detection benchmarks, we\nprovide an ontology for describing anomaly detection contexts, a methodology\nfor controlling various aspects of benchmark creation, guidelines for future\nexperimental design and a discussion of the many potential pitfalls of trying\nto measure success in this field.","url_abs":"http://arxiv.org/abs/1503.01158v2","url_pdf":"http://arxiv.org/pdf/1503.01158v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-meta-analysis-of-the-anomaly-detection","repo_url":"https://github.com/yaroslav-moiseev/evidence-based-possibly-best-practices-in-classical-ML","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"anomaly-detection","task_name":"Anomaly Detection"},{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"experimental-design","task_name":"Experimental Design"},{"task_slug":"hypothesis-testing","task_name":"Two-sample testing"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1503.01158","atlas_url":"https://app.syntology.ai/?focus=1503.01158","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}