{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/advanced-bloom-filter-based-algorithms-for","title":"Advanced Bloom Filter Based Algorithms for Efficient Approximate Data De-Duplication in Streams","arxiv_id":"1212.3964","date":"2012-12-17","proceeding":null,"authors":["Suman K. Bera","Sourav Dutta","Ankur Narang","Souvik Bhattacherjee"],"abstract":"Applications involving telecommunication call data records, web pages, online\ntransactions, medical records, stock markets, climate warning systems, etc.,\nnecessitate efficient management and processing of such massively exponential\namount of data from diverse sources. De-duplication or Intelligent Compression\nin streaming scenarios for approximate identification and elimination of\nduplicates from such unbounded data stream is a greater challenge given the\nreal-time nature of data arrival. Stable Bloom Filters (SBF) addresses this\nproblem to a certain extent. .\n  In this work, we present several novel algorithms for the problem of\napproximate detection of duplicates in data streams. We propose the Reservoir\nSampling based Bloom Filter (RSBF) combining the working principle of reservoir\nsampling and Bloom Filters. We also present variants of the novel Biased\nSampling based Bloom Filter (BSBF) based on biased sampling concepts. We also\npropose a randomized load balanced variant of the sampling Bloom Filter\napproach to efficiently tackle the duplicate detection. In this work, we thus\nprovide a generic framework for de-duplication using Bloom Filters. Using\ndetailed theoretical analysis we prove analytical bounds on the false positive\nrate, false negative rate and convergence rate of the proposed structures. We\nexhibit that our models clearly outperform the existing methods. We also\ndemonstrate empirical analysis of the structures using real-world datasets (3\nmillion records) and also with synthetic datasets (1 billion records) capturing\nvarious input distributions.","url_abs":"http://arxiv.org/abs/1212.3964v1","url_pdf":"http://arxiv.org/pdf/1212.3964v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"advanced-bloom-filter-based-algorithms-for","repo_url":"https://github.com/jeffrey-xiao/probabilistic-collections-rs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"management","task_name":"Management"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}