{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scalable-exact-parent-sets-identification-in","title":"Scalable Exact Parent Sets Identification in Bayesian Networks Learning with Apache Spark","arxiv_id":"1705.06390","date":"2017-05-18","proceeding":null,"authors":["Subhadeep Karan","Jaroslaw Zola"],"abstract":"In Machine Learning, the parent set identification problem is to find a set\nof random variables that best explain selected variable given the data and some\npredefined scoring function. This problem is a critical component to structure\nlearning of Bayesian networks and Markov blankets discovery, and thus has many\npractical applications, ranging from fraud detection to clinical decision\nsupport. In this paper, we introduce a new distributed memory approach to the\nexact parent sets assignment problem. To achieve scalability, we derive\ntheoretical bounds to constraint the search space when MDL scoring function is\nused, and we reorganize the underlying dynamic programming such that the\ncomputational density is increased and fine-grain synchronization is\neliminated. We then design efficient realization of our approach in the Apache\nSpark platform. Through experimental results, we demonstrate that the method\nmaintains strong scalability on a 500-core standalone Spark cluster, and it can\nbe used to efficiently process data sets with 70 variables, far beyond the\nreach of the currently available solutions.","url_abs":"http://arxiv.org/abs/1705.06390v2","url_pdf":"http://arxiv.org/pdf/1705.06390v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scalable-exact-parent-sets-identification-in","repo_url":"https://gitlab.com/SCoRe-Group/SABNA-Release","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"fraud-detection","task_name":"Fraud Detection"}],"methods":[{"method_slug":"mdl","method_name":"MDL"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}