{"url":"/method/smote","slug":"smote","name":"SMOTE","full_name":"Synthetic Minority Over-sampling Technique.","full_name_withheld":false,"description_markdown":"Perhaps the most widely used approach to synthesizing new examples is called the Synthetic Minority Oversampling Technique, or SMOTE for short. This technique was described by Nitesh Chawla, et al. in their 2002 paper named for the technique titled “SMOTE: Synthetic Minority Over-sampling Technique.”\r\n\r\nSMOTE works by selecting examples that are close in the feature space, drawing a line between the examples in the feature space and drawing a new sample at a point along that line.","description_state":"present","introduced_year":null,"introduced_by":{"title":"SMOTE: Synthetic Minority Over-sampling Technique","paper":"/paper/smote-synthetic-minority-over-sampling","first_author":"N. V. Chawla","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/smote-synthetic-minority-over-sampling"},"source":{"url":"http://arxiv.org/abs/1106.1813v1","title":"SMOTE: Synthetic Minority Over-sampling Technique","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Downsampling","url":"/methods/category/downsampling","pwc_aliases":[]}],"n_papers_tagged":156,"archive_num_papers":156,"papers_newest_first":[{"paper":null,"title":"CopulaSMOTE: A Copula-Based Oversampling Approach for Imbalanced Classification in Diabetes Prediction","date":"2025-06-18","arxiv_id":"2506.17326","n_code_links":0,"syntology":null},{"paper":null,"title":"A Comprehensive Analysis of COVID-19 Detection Using Bangladeshi Data and Explainable AI","date":"2025-06-08","arxiv_id":"2506.07234","n_code_links":0,"syntology":null},{"paper":null,"title":"Continuous Fair SMOTE -- Fairness-Aware Stream Learning from Imbalanced Data","date":"2025-05-19","arxiv_id":"2505.13116","n_code_links":0,"syntology":null},{"paper":null,"title":"Data Balancing Strategies: A Survey of Resampling and Augmentation Methods","date":"2025-05-17","arxiv_id":"2505.13518","n_code_links":0,"syntology":null},{"paper":null,"title":"Interactive Diabetes Risk Prediction Using Explainable Machine Learning: A Dash-Based Approach with SHAP, LIME, and Comorbidity Insights","date":"2025-05-08","arxiv_id":"2505.05683","n_code_links":0,"syntology":null},{"paper":null,"title":"VR-FuseNet: A Fusion of Heterogeneous Fundus Data and Explainable Deep Network for Diabetic Retinopathy Classification","date":"2025-04-30","arxiv_id":"2504.21464","n_code_links":0,"syntology":null},{"paper":null,"title":"iHHO-SMOTe: A Cleansed Approach for Handling Outliers and Reducing Noise to Improve Imbalanced Data Classification","date":"2025-04-17","arxiv_id":"2504.12850","n_code_links":0,"syntology":null},{"paper":null,"title":"Kernel-Based Enhanced Oversampling Method for Imbalanced Classification","date":"2025-04-12","arxiv_id":"2504.09147","n_code_links":0,"syntology":null},{"paper":null,"title":"Detecting Credit Card Fraud via Heterogeneous Graph Neural Networks with Graph Attention","date":"2025-04-11","arxiv_id":"2504.08183","n_code_links":0,"syntology":null},{"paper":"/paper/enhancing-metabolic-syndrome-prediction-with","title":"Enhancing Metabolic Syndrome Prediction with Hybrid Data Balancing and Counterfactuals","date":"2025-04-09","arxiv_id":"2504.06987","n_code_links":1,"syntology":null},{"paper":null,"title":"Machine Learning for Identifying Potential Participants in Uruguayan Social Programs","date":"2025-03-31","arxiv_id":"2504.01045","n_code_links":0,"syntology":null},{"paper":null,"title":"Harnessing Mixed Features for Imbalance Data Oversampling: Application to Bank Customers Scoring","date":"2025-03-26","arxiv_id":"2503.22730","n_code_links":0,"syntology":null},{"paper":null,"title":"Optimizing Fire Safety: Reducing False Alarms Using Advanced Machine Learning Techniques","date":"2025-03-13","arxiv_id":"2503.09960","n_code_links":0,"syntology":null},{"paper":null,"title":"ICPR 2024 Competition on Rider Intention Prediction","date":"2025-03-11","arxiv_id":"2503.08437","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards species' classification of the \\textit{Anastrepha pseudoparallela} group","date":"2025-03-11","arxiv_id":"2503.08598","n_code_links":0,"syntology":null},{"paper":null,"title":"Machine learning algorithms to predict stroke in China based on causal inference of time series analysis","date":"2025-03-10","arxiv_id":"2503.14512","n_code_links":0,"syntology":null},{"paper":null,"title":"Attention-Based Synthetic Data Generation for Calibration-Enhanced Survival Analysis: A Case Study for Chronic Kidney Disease Using Electronic Health Records","date":"2025-03-08","arxiv_id":"2503.06096","n_code_links":0,"syntology":null},{"paper":null,"title":"Simplicial SMOTE: Oversampling Solution to the Imbalanced Learning Problem","date":"2025-03-05","arxiv_id":"2503.03418","n_code_links":0,"syntology":null},{"paper":"/paper/llm-tabflow-synthetic-tabular-data-generation","title":"LLM-TabFlow: Synthetic Tabular Data Generation with Inter-column Logical Relationship Preservation","date":"2025-03-04","arxiv_id":"2503.02161","n_code_links":1,"syntology":null},{"paper":null,"title":"Deep learning and classical computer vision techniques in medical image analysis: Case studies on brain MRI tissue segmentation, lung CT COPD registration, and skin lesion classification","date":"2025-02-26","arxiv_id":"2502.19258","n_code_links":0,"syntology":null},{"paper":null,"title":"ML-Driven Approaches to Combat Medicare Fraud: Advances in Class Imbalance Solutions, Feature Engineering, Adaptive Learning, and Business Impact","date":"2025-02-21","arxiv_id":"2502.15898","n_code_links":0,"syntology":null},{"paper":null,"title":"Implementing Large Quantum Boltzmann Machines as Generative AI Models for Dataset Balancing","date":"2025-02-05","arxiv_id":"2502.03086","n_code_links":0,"syntology":null},{"paper":"/paper/assessing-data-augmentation-induced-bias-in","title":"Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models","date":"2025-02-03","arxiv_id":"2502.01825","n_code_links":1,"syntology":null},{"paper":null,"title":"Quantum SMOTE with Angular Outliers: Redefining Minority Class Handling","date":"2025-01-31","arxiv_id":"2501.19001","n_code_links":0,"syntology":null},{"paper":"/paper/a-machine-learning-framework-for-handling","title":"A Machine Learning Framework for Handling Unreliable Absence Label and Class Imbalance for Marine Stinger Beaching Prediction","date":"2025-01-20","arxiv_id":"2501.11293","n_code_links":1,"syntology":null},{"paper":null,"title":"An Imbalanced Learning-based Sampling Method for Physics-informed Neural Networks","date":"2025-01-20","arxiv_id":"2501.11222","n_code_links":0,"syntology":null},{"paper":"/paper/robust-hybrid-classical-quantum-transfer","title":"Robust Hybrid Classical-Quantum Transfer Learning Model for Text Classification Using GPT-Neo 125M with LoRA & SMOTE Enhancement","date":"2025-01-12","arxiv_id":"2501.10435","n_code_links":1,"syntology":null},{"paper":null,"title":"Robust COVID-19 Detection from Cough Sounds using Deep Neural Decision Tree and Forest: A Comprehensive Cross-Datasets Evaluation","date":"2025-01-02","arxiv_id":"2501.01117","n_code_links":0,"syntology":null},{"paper":null,"title":"S&P 500 Trend Prediction","date":"2024-12-16","arxiv_id":"2412.11462","n_code_links":0,"syntology":null},{"paper":null,"title":"Early Diagnosis of Alzheimer's Diseases and Dementia from MRI Images Using an Ensemble Deep Learning","date":"2024-12-07","arxiv_id":"2412.05666","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/classification","name":"General Classification","papers":23},{"task":"/task/classification-1","name":"Classification","papers":22},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":15},{"task":"/task/fraud-detection","name":"Fraud Detection","papers":14},{"task":"/task/feature-selection","name":"feature selection","papers":14},{"task":"/task/imbalanced-classification","name":"imbalanced classification","papers":12},{"task":"/task/intrusion-detection","name":"Intrusion Detection","papers":7},{"task":"/task/diagnostic","name":"Diagnostic","papers":6},{"task":"/task/synthetic-data-generation","name":"Synthetic Data Generation","papers":6},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":5},{"task":"/task/feature-importance","name":"Feature Importance","papers":5},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":5},{"task":"/task/binary-classification","name":"Binary Classification","papers":4},{"task":"/task/deep-learning","name":"Deep Learning","papers":4},{"task":null,"name":"Generative Adversarial Network","papers":4},{"task":"/task/imputation","name":"Imputation","papers":4},{"task":"/task/network-intrusion-detection","name":"Network Intrusion Detection","papers":4},{"task":"/task/text-classification","name":"Text Classification","papers":4},{"task":"/task/time-series-1","name":"Time Series","papers":4},{"task":"/task/time-series","name":"Time Series Analysis","papers":4}],"tasks_shown":20,"n_tasks":115,"usage_by_year":[{"year":"2011","papers":1},{"year":"2014","papers":1},{"year":"2016","papers":2},{"year":"2017","papers":2},{"year":"2018","papers":8},{"year":"2019","papers":11},{"year":"2020","papers":15},{"year":"2021","papers":16},{"year":"2022","papers":26},{"year":"2023","papers":20},{"year":"2024","papers":26},{"year":"2025","papers":28}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/smote"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}