Papers › Scalable Exact Parent Sets Identification in Bayesian Networks Learning with Apache Spark

Scalable Exact Parent Sets Identification in Bayesian Networks Learning with Apache Spark

18 May 2017arXiv:1705.06390archive 2025-07-28

Subhadeep Karan, Jaroslaw Zola

In Machine Learning, the parent set identification problem is to find a set of random variables that best explain selected variable given the data and some predefined scoring function. This problem is a critical component to structure learning of Bayesian networks and Markov blankets discovery, and thus has many practical applications, ranging from fraud detection to clinical decision support. In this paper, we introduce a new distributed memory approach to the exact parent sets assignment problem. To achieve scalability, we derive theoretical bounds to constraint the search space when MDL scoring function is used, and we reorganize the underlying dynamic programming such that the computational density is increased and fine-grain synchronization is eliminated. We then design efficient realization of our approach in the Apache Spark platform. Through experimental results, we demonstrate that the method maintains strong scalability on a 500-core standalone Spark cluster, and it can be used to efficiently process data sets with 70 variables, far beyond the reach of the currently available solutions.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Fraud Detection

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

MDL

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections