{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/is-big-data-sufficient-for-a-reliable","title":"Is Big Data Sufficient for a Reliable Detection of Non-Technical Losses?","arxiv_id":"1702.03767","date":"2017-02-13","proceeding":null,"authors":["Patrick Glauner","Angelo Migliosi","Jorge Meira","Petko Valtchev","Radu State","Franck Bettinger"],"abstract":"Non-technical losses (NTL) occur during the distribution of electricity in\npower grids and include, but are not limited to, electricity theft and faulty\nmeters. In emerging countries, they may range up to 40% of the total\nelectricity distributed. In order to detect NTLs, machine learning methods are\nused that learn irregular consumption patterns from customer data and\ninspection results. The Big Data paradigm followed in modern machine learning\nreflects the desire of deriving better conclusions from simply analyzing more\ndata, without the necessity of looking at theory and models. However, the\nsample of inspected customers may be biased, i.e. it does not represent the\npopulation of all customers. As a consequence, machine learning models trained\non these inspection results are biased as well and therefore lead to unreliable\npredictions of whether customers cause NTL or not. In machine learning, this\nissue is called covariate shift and has not been addressed in the literature on\nNTL detection yet. In this work, we present a novel framework for quantifying\nand visualizing covariate shift. We apply it to a commercial data set from\nBrazil that consists of 3.6M customers and 820K inspection results. We show\nthat some features have a stronger covariate shift than others, making\npredictions less reliable. In particular, previous inspections were focused on\ncertain neighborhoods or customer classes and that they were not sufficiently\nspread among the population of customers. This framework is about to be\ndeployed in a commercial product for NTL detection.","url_abs":"http://arxiv.org/abs/1702.03767v2","url_pdf":"http://arxiv.org/pdf/1702.03767v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"is-big-data-sufficient-for-a-reliable","repo_url":"https://github.com/pglauner/SpatialBiasNTL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}