{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/expected-similarity-estimation-for-large","title":"Expected Similarity Estimation for Large-Scale Batch and Streaming Anomaly Detection","arxiv_id":"1601.06602","date":"2016-01-25","proceeding":null,"authors":["Markus Schneider","Wolfgang Ertel","Fabio Ramos"],"abstract":"We present a novel algorithm for anomaly detection on very large datasets and\ndata streams. The method, named EXPected Similarity Estimation (EXPoSE), is\nkernel-based and able to efficiently compute the similarity between new data\npoints and the distribution of regular data. The estimator is formulated as an\ninner product with a reproducing kernel Hilbert space embedding and makes no\nassumption about the type or shape of the underlying data distribution. We show\nthat offline (batch) learning with EXPoSE can be done in linear time and online\n(incremental) learning takes constant time per instance and model update.\nFurthermore, EXPoSE can make predictions in constant time, while it requires\nonly constant memory. In addition, we propose different methodologies for\nconcept drift adaptation on evolving data streams. On several real datasets we\ndemonstrate that our approach can compete with state of the art algorithms for\nanomaly detection while being an order of magnitude faster than most other\napproaches.","url_abs":"http://arxiv.org/abs/1601.06602v3","url_pdf":"http://arxiv.org/pdf/1601.06602v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"expected-similarity-estimation-for-large","repo_url":"https://github.com/numenta/NAB","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"anomaly-detection","task_name":"Anomaly Detection"},{"task_slug":"incremental-learning","task_name":"Incremental Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}