{"url":"/method/tabnet","slug":"tabnet","name":"TabNet","full_name":"TabNet","full_name_withheld":false,"description_markdown":"**TabNet** is a deep tabular data learning architecture that uses sequential attention to choose which features to reason from at each decision step.\r\n\r\nThe TabNet encoder is composed of a feature transformer, an attentive transformer and feature masking. A split block\r\ndivides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the masks can be aggregated to obtain global feature important attribution. The TabNet decoder is composed of a feature transformer block at each step. \r\n\r\nIn the feature transformer block, a 4-layer network is used, where 2 are shared across all decision steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates how much each feature has been used before the current decision step. sparsemax is used for normalization of the coefficients, resulting in sparse selection of the salient features.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1908.07442v5","title":"TabNet: Attentive Interpretable Tabular Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Deep Tabular Learning","url":"/methods/category/deep-tabular-learning","pwc_aliases":[]}],"n_papers_tagged":29,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"The Impact of Feature Scaling In Machine Learning: Effects on Regression and Classification Tasks","date":"2025-06-09","arxiv_id":"2506.08274","n_code_links":0,"syntology":null},{"paper":null,"title":"DuAL-Net: A Hybrid Framework for Alzheimer's Disease Prediction from Whole-Genome Sequencing via Local SNP Windows and Global Annotations","date":"2025-05-31","arxiv_id":"2506.00673","n_code_links":0,"syntology":null},{"paper":null,"title":"Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification","date":"2025-05-22","arxiv_id":"2505.16338","n_code_links":0,"syntology":null},{"paper":"/paper/benchmarking-traditional-machine-learning-and","title":"Benchmarking Traditional Machine Learning and Deep Learning Models for Fault Detection in Power Transformers","date":"2025-05-07","arxiv_id":"2505.06295","n_code_links":1,"syntology":null},{"paper":null,"title":"Can Moran Eigenvectors Improve Machine Learning of Spatial Data? Insights from Synthetic Data Validation","date":"2025-04-16","arxiv_id":"2504.12450","n_code_links":0,"syntology":null},{"paper":"/paper/enhancing-metabolic-syndrome-prediction-with","title":"Enhancing Metabolic Syndrome Prediction with Hybrid Data Balancing and Counterfactuals","date":"2025-04-09","arxiv_id":"2504.06987","n_code_links":1,"syntology":null},{"paper":null,"title":"A Survey on Deep Tabular Learning","date":"2024-10-15","arxiv_id":"2410.12034","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhanced Credit Score Prediction Using Ensemble Deep Learning Model","date":"2024-09-30","arxiv_id":"2410.00256","n_code_links":0,"syntology":null},{"paper":"/paper/gradient-boosting-decision-trees-on-medical","title":"Gradient Boosting Decision Trees on Medical Diagnosis over Tabular Data","date":"2024-09-25","arxiv_id":"2410.03705","n_code_links":1,"syntology":null},{"paper":null,"title":"Advancing Machine Learning in Industry 4.0: Benchmark Framework for Rare-event Prediction in Chemical Processes","date":"2024-08-31","arxiv_id":"2409.00485","n_code_links":0,"syntology":null},{"paper":"/paper/interpretable-graph-neural-networks-for-2","title":"Interpretable Graph Neural Networks for Heterogeneous Tabular Data","date":"2024-08-14","arxiv_id":"2408.07661","n_code_links":2,"syntology":{"ran":4,"of":5,"unverified":1,"pointer_only":0}},{"paper":"/paper/interpretabnet-distilling-predictive-signals","title":"InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature Interpretation","date":"2024-06-01","arxiv_id":"2406.00426","n_code_links":1,"syntology":{"ran":10,"of":15,"unverified":5,"pointer_only":0}},{"paper":"/paper/is-interpretable-machine-learning-effective","title":"Is Interpretable Machine Learning Effective at Feature Selection for Neural Learning-to-Rank?","date":"2024-05-13","arxiv_id":"2405.07782","n_code_links":1,"syntology":null},{"paper":null,"title":"Federated Learning for Tabular Data using TabNet: A Vehicular Use-Case","date":"2024-05-03","arxiv_id":"2405.02060","n_code_links":0,"syntology":null},{"paper":null,"title":"TabVFL: Improving Latent Representation in Vertical Federated Learning","date":"2024-04-27","arxiv_id":"2404.17990","n_code_links":0,"syntology":null},{"paper":null,"title":"FH-TabNet: Multi-Class Familial Hypercholesterolemia Detection via a Multi-Stage Tabular Deep Learning","date":"2024-03-16","arxiv_id":"2403.11032","n_code_links":0,"syntology":null},{"paper":"/paper/exploring-factors-affecting-pedestrian-crash","title":"Exploring Factors Affecting Pedestrian Crash Severity Using TabNet: A Deep Learning Approach","date":"2023-11-29","arxiv_id":"2312.00066","n_code_links":1,"syntology":null},{"paper":null,"title":"Classification Methods Based on Machine Learning for the Analysis of Fetal Health Data","date":"2023-11-18","arxiv_id":"2311.10962","n_code_links":0,"syntology":null},{"paper":null,"title":"Stable and Interpretable Deep Learning for Tabular Data: Introducing InterpreTabNet with the Novel InterpreStability Metric","date":"2023-10-04","arxiv_id":"2310.02870","n_code_links":0,"syntology":null},{"paper":"/paper/interpretable-graph-neural-networks-for-1","title":"Interpretable Graph Neural Networks for Tabular Data","date":"2023-08-17","arxiv_id":"2308.08945","n_code_links":2,"syntology":null},{"paper":"/paper/machine-learning-methods-for-the-search-for-l","title":"Machine learning methods for the search for L&T brown dwarfs in the data of modern sky surveys","date":"2023-08-06","arxiv_id":"2308.03045","n_code_links":1,"syntology":null},{"paper":null,"title":"Pump It Up: Predict Water Pump Status using Attentive Tabular Learning","date":"2023-04-08","arxiv_id":"2304.03969","n_code_links":0,"syntology":null},{"paper":null,"title":"Tiny Classifier Circuits: Evolving Accelerators for Tabular Data","date":"2023-02-28","arxiv_id":"2303.00031","n_code_links":0,"syntology":null},{"paper":null,"title":"Interpreting Black-box Machine Learning Models for High Dimensional Datasets","date":"2022-08-29","arxiv_id":"2208.13405","n_code_links":0,"syntology":null},{"paper":null,"title":"Don't read, just look: Main content extraction from web pages using visual features","date":"2021-10-27","arxiv_id":"2110.14164","n_code_links":0,"syntology":null},{"paper":null,"title":"Analysis of Vision-based Abnormal Red Blood Cell Classification","date":"2021-06-01","arxiv_id":"2106.00389","n_code_links":0,"syntology":null},{"paper":"/paper/pytorch-tabular-a-framework-for-deep-learning","title":"PyTorch Tabular: A Framework for Deep Learning with Tabular Data","date":"2021-04-28","arxiv_id":"2104.13638","n_code_links":1,"syntology":null},{"paper":null,"title":"Fairness in TabNet Model by Disentangled Representation for the Prediction of Hospital No-Show","date":"2021-03-06","arxiv_id":"2103.04048","n_code_links":0,"syntology":null},{"paper":"/paper/tabnet-attentive-interpretable-tabular","title":"TabNet: Attentive Interpretable Tabular Learning","date":"2019-08-20","arxiv_id":"1908.07442","n_code_links":19,"syntology":{"ran":1,"of":17,"unverified":16,"pointer_only":1}}],"papers_shown":29,"tasks":[{"task":"/task/representation-learning","name":"Representation Learning","papers":4},{"task":"/task/feature-selection","name":"feature selection","papers":4},{"task":"/task/classification-1","name":"Classification","papers":3},{"task":"/task/decision-making","name":"Decision Making","papers":2},{"task":"/task/deep-learning","name":"Deep Learning","papers":2},{"task":"/task/federated-learning","name":"Federated Learning","papers":2},{"task":"/task/graph-neural-network","name":"Graph Neural Network","papers":2},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":1},{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":1},{"task":"/task/benchmarking","name":"Benchmarking","papers":1},{"task":"/task/binary-classification","name":"Binary Classification","papers":1},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":1},{"task":"/task/credit-score","name":"Credit score","papers":1},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/dimensionality-reduction","name":"Dimensionality Reduction","papers":1},{"task":"/task/disease-prediction","name":"Disease Prediction","papers":1},{"task":"/task/edge-computing","name":"Edge-computing","papers":1},{"task":"/task/explainable-models","name":"Explainable Models","papers":1},{"task":"/task/fairness","name":"Fairness","papers":1},{"task":"/task/fault-detection","name":"Fault Detection","papers":1}],"tasks_shown":20,"n_tasks":49,"usage_by_year":[{"year":"2019","papers":1},{"year":"2021","papers":4},{"year":"2022","papers":1},{"year":"2023","papers":7},{"year":"2024","papers":10},{"year":"2025","papers":6}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/tabnet"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}