{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cusboost-cluster-based-under-sampling-with","title":"CUSBoost: Cluster-based Under-sampling with Boosting for Imbalanced Classification","arxiv_id":"1712.04356","date":"2017-12-12","proceeding":null,"authors":["Farshid Rayhan","Sajid Ahmed","Asif Mahbub","Md. Rafsan Jani","Swakkhar Shatabda","Dewan Md. Farid"],"abstract":"Class imbalance classification is a challenging research problem in data\nmining and machine learning, as most of the real-life datasets are often\nimbalanced in nature. Existing learning algorithms maximise the classification\naccuracy by correctly classifying the majority class, but misclassify the\nminority class. However, the minority class instances are representing the\nconcept with greater interest than the majority class instances in real-life\napplications. Recently, several techniques based on sampling methods\n(under-sampling of the majority class and over-sampling the minority class),\ncost-sensitive learning methods, and ensemble learning have been used in the\nliterature for classifying imbalanced datasets. In this paper, we introduce a\nnew clustering-based under-sampling approach with boosting (AdaBoost)\nalgorithm, called CUSBoost, for effective imbalanced classification. The\nproposed algorithm provides an alternative to RUSBoost (random under-sampling\nwith AdaBoost) and SMOTEBoost (synthetic minority over-sampling with AdaBoost)\nalgorithms. We evaluated the performance of CUSBoost algorithm with the\nstate-of-the-art methods based on ensemble learning like AdaBoost, RUSBoost,\nSMOTEBoost on 13 imbalance binary and multi-class datasets with various\nimbalance ratios. The experimental results show that the CUSBoost is a\npromising and effective approach for dealing with highly imbalanced datasets.","url_abs":"http://arxiv.org/abs/1712.04356v1","url_pdf":"http://arxiv.org/pdf/1712.04356v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cusboost-cluster-based-under-sampling-with","repo_url":"https://github.com/farshidrayhanuiu/CUSBoost","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"ensemble-learning","task_name":"Ensemble Learning"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"imbalanced-classification","task_name":"imbalanced classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.04356","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}