{"url":"/dataset/pubchem18","name":"PubChem18","full_name":"PubChem 2018","description_markdown":"A.2.1 AN OPEN, LARGE-SCALE DATASET FOR ZERO-SHOT DRUG DISCOVERY DERIVED FROM PUBCHEM\r\nWe constructed a large public dataset extracted from PubChem (Kim et al., 2019; Preuer et al., 2018), an open chemistry\r\ndatabase, and the largest collection of readily available chemical data. We take assays ranging from 2004 to 2018-05.\r\nIt initially comprises 224,290,250 records of molecule-bioassay activity, corresponding to 2,120,854 unique molecules\r\nand 21,003 unique bioassays. We find that some molecule-bioassay pairs have multiple activity records, which may not\r\nall agree. We reduce every molecule-bioassay pair to exactly one activity measurement by applying majority voting.\r\nMolecule-bioassay pairs with ties are discarded. This step yields our final bioactivity dataset, which features 223,219,241 records of molecule-bioassay activity, corresponding to 2,120,811 unique molecules and 21,002 unique bioassays ranging\r\nfrom AID 1 to AID 1259411. Molecules range up to CID 132472079. The dataset has 3 different splitting schemes.","description_withheld":null,"homepage":"","introduced_date":"2023-03-06","introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Biology","url":"/datasets/modality/biology"}],"tasks":[{"name":"Zero-Shot Learning","url":"/task/zero-shot-learning","datasets_with_task":"/datasets/task/zero-shot-learning"},{"name":"Molecular Property Prediction","url":"/task/molecular-property-prediction","datasets_with_task":"/datasets/task/molecular-property-prediction"},{"name":"Molecular Property Prediction (1-shot))","url":"/task/molecular-property-prediction-1-shot","datasets_with_task":"/datasets/task/molecular-property-prediction-1-shot"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["PubChem18"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}