Papers › Curvature-based Feature Selection with Application in Classifying Electronic Health Records

Curvature-based Feature Selection with Application in Classifying Electronic Health Records

10 Jan 2021arXiv:2101.03581archive 2025-07-28

Zheming Zuo, Jie Li, Han Xu, Noura Al Moubayed

Disruptive technologies provides unparalleled opportunities to contribute to the identifications of many aspects in pervasive healthcare, from the adoption of the Internet of Things through to Machine Learning (ML) techniques. As a powerful tool, ML has been widely applied in patient-centric healthcare solutions. To further improve the quality of patient care, Electronic Health Records (EHRs) are commonly adopted in healthcare facilities for analysis. It is a crucial task to apply AI and ML to analyse those EHRs for prediction and diagnostics due to their highly unstructured, unbalanced, incomplete, and high-dimensional nature. Dimensionality reduction is a common data preprocessing technique to cope with high-dimensional EHR data, which aims to reduce the number of features of EHR representation while improving the performance of the subsequent data analysis, e.g. classification. In this work, an efficient filter-based feature selection method, namely Curvature-based Feature Selection (CFS), is presented. The proposed CFS applied the concept of Menger Curvature to rank the weights of all features in the given data set. The performance of the proposed CFS has been evaluated in four well-known EHR data sets, including Cervical Cancer Risk Factors (CCRFDS), Breast Cancer Coimbra (BCCDS), Breast Tissue (BTDS), and Diabetic Retinopathy Debrecen (DRDDS). The experimental results show that the proposed CFS achieved state-of-the-art performance on the above data sets against conventional PCA and other most recent approaches. The source code of the proposed approach is publicly available at https://github.com/zhemingzuo/CFS.

PaperPDFCode

Code

zhemingzuo/CFS officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Breast Cancer DetectionBreast Tissue IdentificationCervical cancer biopsy identificationDiabetic Retinopathy DetectionDimensionality Reductionfeature selection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Breast Cancer Detection Breast Cancer Coimbra Data Set CFS-TSK+ Mean Accuracy 79.17 #1 of 1 Archive leaderboard report
Breast Tissue Identification Breast Tissue Data Set CFS-QDA Mean Accuracy 100.00 #1 of 1 Archive leaderboard report
Cervical cancer biopsy identification Cervical Cancer (Risk Factors) Data Set CFS-TSK+ Mean Accuracy 97.09 #1 of 1 Archive leaderboard report
Diabetic Retinopathy Detection Diabetic Retinopathy Debrecen Data Set CFS-BPNN Mean Accuracy 74.72 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Feature SelectionPCA

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections