{"url":"/dataset/dermatology-ddx-dataset","name":"Dermatology ddx dataset","full_name":null,"description_markdown":"The dermatology differential diagnoses (ddx) dataset for skin condition classification includes expert annotations and model predictions for 1947 cases. Note that no images or meta information are provided. The expert annotations come in the form of differential diagnoses, i.e., partial rankings of conditions, and there is a high level of disagreement among experts, making this a perfect benchmark for dealing with disagreement. The data has been introduced in [[1]](https://arxiv.org/abs/2307.02191) and [[2]](https://arxiv.org/abs/2307.09302).\r\n\r\n```\r\n[1] Stutz, D., Roy, A.G., Matejovicova, T., Strachan, P., Cemgil, A.T.,\r\n    & Doucet, A. (2023).\r\n    [Conformal prediction under ambiguous ground truth](https://openreview.net/forum?id=CAd6V2qXxc).\r\n    TMLR.\r\n[2] Stutz, D., Cemgil, A.T., Roy, A.G., Matejovicova, T., Barsbey, M., Strachan,\r\n   P., Schaekermann, M., Freyberg, J.V., Rikhye, R.V., Freeman, B., Matos, J.P.,\r\n   Telang, U., Webster, D.R., Liu, Y., Corrado, G.S., Matias, Y., Kohli, P.,\r\n   Liu, Y., Doucet, A., & Karthikesalingam, A. (2023).\r\n   [Evaluating AI systems under uncertain ground truth: a case study in dermatology](https://arxiv.org/abs/2307.02191).\r\n   ArXiv, abs/2307.02191.\r\n```\r\n\r\nThe dataset is structured as follows:\r\n\r\n*  `data/dermatology_selectors.json`: The expert annotations as partial\r\n   rankings. These partial rankings are encoded as so-called \"selectors\":\r\n   For each case, there are multiple partial rankings, each partial\r\n   ranking is a list of grouped classes (i.e., skin conditions).\r\n   Teh example from [1, 2] shown below describes a partial ranking where\r\n   \"Hemangioma\" is ranked first followed by a group of three conditions,\r\n   including \"Melanocytic Nevus\", \"Melanoma\", and \"O/E\". In the JSON file,\r\n   the conditions are encoded as numbers and the mapping of numbers to\r\n   condition names can be found in `data/dermatology_conditions.txt`.\r\n\r\n```\r\n['Hemangioma'], ['Melanocytic Nevus', 'Melanoma', 'O/E']\r\n```\r\n\r\n*  `data/dermatology_predictions[0-4].json`: Model predictions of models\r\n   A to D in [1] as `1947 x 419` float arrays saved using `numpy.savetxt` with\r\n   `fmt='%.3e'`.\r\n*  `data/dermatology_conditions.txt`: Condition names for each class.\r\n*  `data/dermatology_risks.txt`: Risk category for each condition, where\r\n   0 corresponds to low risk, 1 to medium risk and 2 to high risk.","description_withheld":null,"homepage":"https://github.com/google-deepmind/uncertain_ground_truth","introduced_date":"2024-03-28","introduced_date_note":null,"introduced_by":{"paper":"/paper/conformal-prediction-under-ambiguous-ground","title":"Conformal prediction under ambiguous ground truth","first_author":"David Stutz","url":null},"license":{"name":"Creative Commons Attribution 4.0 International License (CC-BY)","url":"https://github.com/google-deepmind/uncertain_ground_truth"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Dermatology ddx dataset"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}