{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/continuous-outlier-mining-of-streaming-data","title":"Continuous Outlier Mining of Streaming Data in Flink","arxiv_id":"1902.07901","date":"2019-02-21","proceeding":null,"authors":["Theodoros Toliopoulos","Anastasios Gounaris","Kostas Tsichlas","Apostolos Papadopoulos","Sandra Sampaio"],"abstract":"In this work, we focus on distance-based outliers in a metric space, where\nthe status of an entity as to whether it is an outlier is based on the number\nof other entities in its neighborhood. In recent years, several solutions have\ntackled the problem of distance-based outliers in data streams, where outliers\nmust be mined continuously as new elements become available. An interesting\nresearch problem is to combine the streaming environment with massively\nparallel systems to provide scalable streambased algorithms. However, none of\nthe previously proposed techniques refer to a massively parallel setting. Our\nproposal fills this gap and investigates the challenges in transferring\nstate-of-the-art techniques to Apache Flink, a modern platform for intensive\nstreaming analytics. We thoroughly present the technical challenges encountered\nand the alternatives that may be applied. We show speed-ups of up to 117 (resp.\n2076) times over a naive parallel (resp. non-parallel) solution in Flink, by\nusing just an ordinary four-core machine and a real-world dataset. When moving\nto a three-machine cluster, due to less contention, we manage to achieve both\nbetter scalability in terms of the window slide size and the data\ndimensionality, and even higher speed-ups, e.g., by a factor of 510. Overall,\nour results demonstrate that oulier mining can be achieved in an efficient and\nscalable manner. The resulting techniques have been made publicly available as\nopen-source software.","url_abs":"http://arxiv.org/abs/1902.07901v1","url_pdf":"http://arxiv.org/pdf/1902.07901v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"continuous-outlier-mining-of-streaming-data","repo_url":"https://github.com/tatoliop/parallel-streaming-outlier-detection","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}