{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/feature-selection-a-data-perspective","title":"Feature Selection: A Data Perspective","arxiv_id":"1601.07996","date":"2016-01-29","proceeding":null,"authors":["Jundong Li","Kewei Cheng","Suhang Wang","Fred Morstatter","Robert P. Trevino","Jiliang Tang","Huan Liu"],"abstract":"Feature selection, as a data preprocessing strategy, has been proven to be\neffective and efficient in preparing data (especially high-dimensional data)\nfor various data mining and machine learning problems. The objectives of\nfeature selection include: building simpler and more comprehensible models,\nimproving data mining performance, and preparing clean, understandable data.\nThe recent proliferation of big data has presented some substantial challenges\nand opportunities to feature selection. In this survey, we provide a\ncomprehensive and structured overview of recent advances in feature selection\nresearch. Motivated by current challenges and opportunities in the era of big\ndata, we revisit feature selection research from a data perspective and review\nrepresentative feature selection algorithms for conventional data, structured\ndata, heterogeneous data and streaming data. Methodologically, to emphasize the\ndifferences and similarities of most existing feature selection algorithms for\nconventional data, we categorize them into four main groups: similarity based,\ninformation theoretical based, sparse learning based and statistical based\nmethods. To facilitate and promote the research in this community, we also\npresent an open-source feature selection repository that consists of most of\nthe popular feature selection algorithms\n(\\url{http://featureselection.asu.edu/}). Also, we use it as an example to show\nhow to evaluate feature selection algorithms. At the end of the survey, we\npresent a discussion about some open problems and challenges that require more\nattention in future research.","url_abs":"http://arxiv.org/abs/1601.07996v5","url_pdf":"http://arxiv.org/pdf/1601.07996v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"feature-selection-a-data-perspective","repo_url":"https://github.com/jundongl/scikit-feature","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"feature-selection-a-data-perspective","repo_url":"https://github.com/TFG-Informatica-Enfermedad-Celiaca/Analisis-EC","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"sparse-learning","task_name":"Sparse Learning"},{"task_slug":"survey","task_name":"Survey"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[],"datasets_introduced":[{"slug":"pixraw10p","name":"pixraw10P","full_name":"pixraw10P"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}