{"url":"/dataset/diabetes","name":"Diabetes","full_name":"Diabetes 130-US Hospitals for Years 1999-2008","description_markdown":"**What do the instances in this dataset represent?**\r\n\r\nThe instances represent hospitalized patient records diagnosed with diabetes.\r\n\r\n**Are there recommended data splits?**\r\n\r\nNo recommendation. The standard train-test split could be used. Can use three-way holdout split (i.e., train-validation-test) when doing model selection.\r\n\r\n**Does the dataset contain data that might be considered sensitive in any way?**\r\n\r\nYes. The dataset contains information about the age, gender, and race of the patients.\r\n\r\n**Additional Information**\r\n\r\nThe dataset represents ten years (1999-2008) of clinical care at 130 US hospitals and integrated delivery networks. It includes over 50 features representing patient and hospital outcomes. Information was extracted from the database for encounters that satisfied the following criteria.\r\n(1)\tIt is an inpatient encounter (a hospital admission).\r\n(2)\tIt is a diabetic encounter, that is, one during which any kind of diabetes was entered into the system as a diagnosis.\r\n(3)\tThe length of stay was at least 1 day and at most 14 days.\r\n(4)\tLaboratory tests were performed during the encounter.\r\n(5)\tMedications were administered during the encounter.\r\n\r\nThe data contains such attributes as patient number, race, gender, age, admission type, time in hospital, medical specialty of admitting physician, number of lab tests performed, HbA1c test result, diagnosis, number of medications, diabetic medications, number of outpatient, inpatient, and emergency visits in the year before the hospitalization, etc.\r\n\r\n**Has Missing Values?**\r\n\r\nYes","description_withheld":null,"homepage":"https://archive.ics.uci.edu/dataset/296/diabetes+130-us+hospitals+for+years+1999-2008","introduced_date":"2014-05-02","introduced_date_note":null,"introduced_by":null,"license":{"name":"CC BY 4.0","url":null},"modalities":[{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Feature Importance","url":"/task/feature-importance","datasets_with_task":"/datasets/task/feature-importance"},{"name":"Tabular Data Generation","url":"/task/tabular-data-generation","datasets_with_task":"/datasets/task/tabular-data-generation"},{"name":"Diabetes Prediction","url":"/task/diabetes-prediction","datasets_with_task":"/datasets/task/diabetes-prediction"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Diabetes"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/tabular-data-generation-on-diabetes","task":"Tabular Data Generation","dataset_variant":"Diabetes","rows":6,"metrics":["DT Accuracy","Parameters(M)","LR Accuracy","RF Accuracy"],"first_row_in_archive_order":{"model":"Binary Diffusion","paper":"/paper/tabular-data-generation-using-binary","metrics":{"DT Accuracy":"0.5713","LR Accuracy":"0.5775","Parameters(M)":"1.8","RF Accuracy":"0.5752"},"code_links":[{"title":"vkinakh/binary-diffusion-tabular","url":"https://github.com/vkinakh/binary-diffusion-tabular"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/feature-importance-on-diabetes","task":"Feature Importance","dataset_variant":"Diabetes","rows":2,"metrics":["Pearson Correlation"],"first_row_in_archive_order":{"model":"VarImpVIANN","paper":"/paper/variance-based-feature-importance-in-neural","metrics":{"Pearson Correlation":"0.86"},"code_links":[{"title":"rebelosa/feature-importance-neural-networks","url":"https://github.com/rebelosa/feature-importance-neural-networks"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/diabetes-prediction-on-diabetes","task":"Diabetes Prediction","dataset_variant":"Diabetes","rows":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"XBNET","paper":"/paper/xbnet-an-extremely-boosted-neural-network","metrics":{"Accuracy":"78.78"},"code_links":[{"title":"tusharsarkar3/XBNet","url":"https://github.com/tusharsarkar3/XBNet"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/tabular-data-generation-using-binary","title":"Tabular Data Generation using Binary Diffusion","date":"2024-09-20","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":10,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/language-models-are-realistic-tabular-data","title":"Language Models are Realistic Tabular Data Generators","date":"2022-10-12","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/xbnet-an-extremely-boosted-neural-network","title":"XBNet : An Extremely Boosted Neural Network","date":"2021-06-09","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/variance-based-feature-importance-in-neural","title":"Variance-Based Feature Importance in Neural Networks","date":"2019-10-16","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/modeling-tabular-data-using-conditional-gan","title":"Modeling Tabular data using Conditional GAN","date":"2019-07-01","rows_on_this_dataset":3,"code_links":9,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":19,"samples_ran":12,"samples_unverified":7,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}