{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cfm-bd-a-distributed-rule-induction-algorithm","title":"CFM-BD: a distributed rule induction algorithm for building Compact Fuzzy Models in Big Data classification problems","arxiv_id":"1902.09357","date":"2019-02-25","proceeding":null,"authors":["Mikel Elkano","Jose Sanz","Edurne Barrenechea","Humberto Bustince","Mikel Galar"],"abstract":"Interpretability has always been a major concern for fuzzy rule-based\nclassifiers. The usage of human-readable models allows them to explain the\nreasoning behind their predictions and decisions. However, when it comes to Big\nData classification problems, fuzzy rule-based classifiers have not been able\nto maintain the good trade-off between accuracy and interpretability that has\ncharacterized these techniques in non-Big Data environments. The most accurate\nmethods build too complex models composed of a large number of rules and fuzzy\nsets, while those approaches focusing on interpretability do not provide\nstate-of-the-art discrimination capabilities. In this paper, we propose a new\ndistributed learning algorithm named CFM-BD to construct accurate and compact\nfuzzy rule-based classification systems for Big Data. This method has been\nspecifically designed from scratch for Big Data problems and does not adapt or\nextend any existing algorithm. The proposed learning process consists of three\nstages: 1) pre-processing based on the probability integral transform theorem;\n2) rule induction inspired by CHI-BD and Apriori algorithms; 3) rule selection\nby means of a global evolutionary optimization. We conducted a complete\nempirical study to test the performance of our approach in terms of accuracy,\ncomplexity, and runtime. The results obtained were compared and contrasted with\nfour state-of-the-art fuzzy classifiers for Big Data (FBDT, FMDT, Chi-Spark-RS,\nand CHI-BD). According to this study, CFM-BD is able to provide competitive\ndiscrimination capabilities using significantly simpler models composed of a\nfew rules of less than 3 antecedents, employing 5 linguistic labels for all\nvariables.","url_abs":"http://arxiv.org/abs/1902.09357v1","url_pdf":"http://arxiv.org/pdf/1902.09357v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cfm-bd-a-distributed-rule-induction-algorithm","repo_url":"https://github.com/melkano/cfm-bd","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}