{"url":"/dataset/pd4ml","name":"pd4ml","full_name":"Physics Data for Machine Learning","description_markdown":"**pd4ml** is a collection of datasets from fundamental physics research -- including particle physics, astroparticle physics, and hadron- and nuclear physics -- for supervised machine learning studies. These datasets, containing hadronic top quarks, cosmic-ray induced air showers, phase transitions in hadronic matter, and generator-level histories, are made public to simplify future work on cross-disciplinary machine learning and transfer learning in fundamental physics.\r\n\r\nIt currently consists on 5 datasets:\r\n\r\n- Top Tagging Landscape (Classification) \r\n    - Train/val/test: 1.2M/400k/400k \r\n    - Structure: Four vectors \r\n    - Dimension: 200 particles, 4 features/particle\r\n- Smart Backgrounds (Classification)   \r\n    - Train/val/test: 157k/39k/84k\r\n    - Structure: Decay Graph\r\n    - Dimension: 100 particles, 9 features/particle\r\n- Spinodal or Not (Classification)   \r\n    - Train/val/test: 16.3k/4k/8.7k\r\n    - Structure: 2D Histogram\r\n    - Dimension: 20x20 histogram of pion spectra\r\n- EoS (Classification)  \r\n    - Train/val/test: 121k/25k/54k\r\n    - Structure: 2D Histogram\r\n    - Dimension: 24x24 histogram of pion spectra\r\n- Air Showers (Regression)   \r\n    - Train/val/test: 56k/30k/14k\r\n    - Structure: 81 1D Traces\r\n    - Dimension: 81 stations, 80 signal bins + timing","description_withheld":null,"homepage":"https://github.com/erum-data-idt/pd4ml","introduced_date":"2021-07-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/shared-data-and-algorithms-for-deep-learning","title":"Shared Data and Algorithms for Deep Learning in Fundamental Physics","first_author":"Lisa Benato","url":null},"license":null,"modalities":[{"name":"Physics","url":"/datasets/modality/physics"}],"tasks":[],"languages":[],"variants":["pd4ml"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}