{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-relief-based-feature-selection","title":"Benchmarking Relief-Based Feature Selection Methods for Bioinformatics Data Mining","arxiv_id":"1711.08477","date":"2017-11-22","proceeding":null,"authors":["Ryan J. Urbanowicz","Randal S. Olson","Peter Schmitt","Melissa Meeker","Jason H. Moore"],"abstract":"Modern biomedical data mining requires feature selection methods that can (1)\nbe applied to large scale feature spaces (e.g. `omics' data), (2) function in\nnoisy problems, (3) detect complex patterns of association (e.g. gene-gene\ninteractions), (4) be flexibly adapted to various problem domains and data\ntypes (e.g. genetic variants, gene expression, and clinical data) and (5) are\ncomputationally tractable. To that end, this work examines a set of\nfilter-style feature selection algorithms inspired by the `Relief' algorithm,\ni.e. Relief-Based algorithms (RBAs). We implement and expand these RBAs in an\nopen source framework called ReBATE (Relief-Based Algorithm Training\nEnvironment). We apply a comprehensive genetic simulation study comparing\nexisting RBAs, a proposed RBA called MultiSURF, and other established feature\nselection methods, over a variety of problems. The results of this study (1)\nsupport the assertion that RBAs are particularly flexible, efficient, and\npowerful feature selection methods that differentiate relevant features having\nunivariate, multivariate, epistatic, or heterogeneous associations, (2) confirm\nthe efficacy of expansions for classification vs. regression, discrete vs.\ncontinuous features, missing data, multiple classes, or class imbalance, (3)\nidentify previously unknown limitations of specific RBAs, and (4) suggest that\nwhile MultiSURF* performs best for explicitly identifying pure 2-way\ninteractions, MultiSURF yields the most reliable feature selection performance\nacross a wide range of problem types.","url_abs":"http://arxiv.org/abs/1711.08477v2","url_pdf":"http://arxiv.org/pdf/1711.08477v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-relief-based-feature-selection","repo_url":"https://github.com/EpistasisLab/ReBATE","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"benchmarking-relief-based-feature-selection","repo_url":"https://github.com/EpistasisLab/scikit-rebate","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}