{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/therapeutics-data-commons-machine-learning","title":"Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development","arxiv_id":"2102.09548","date":"2021-02-18","proceeding":null,"authors":["Kexin Huang","Tianfan Fu","Wenhao Gao","Yue Zhao","Yusuf Roohani","Jure Leskovec","Connor W. Coley","Cao Xiao","Jimeng Sun","Marinka Zitnik"],"abstract":"Therapeutics machine learning is an emerging field with incredible opportunities for innovatiaon and impact. However, advancement in this field requires formulation of meaningful learning tasks and careful curation of datasets. Here, we introduce Therapeutics Data Commons (TDC), the first unifying platform to systematically access and evaluate machine learning across the entire range of therapeutics. To date, TDC includes 66 AI-ready datasets spread across 22 learning tasks and spanning the discovery and development of safe and effective medicines. TDC also provides an ecosystem of tools and community resources, including 33 data functions and types of meaningful data splits, 23 strategies for systematic model evaluation, 17 molecule generation oracles, and 29 public leaderboards. All resources are integrated and accessible via an open Python library. We carry out extensive experiments on selected datasets, demonstrating that even the strongest algorithms fall short of solving key therapeutics challenges, including real dataset distributional shifts, multi-scale modeling of heterogeneous data, and robust generalization to novel data points. We envision that TDC can facilitate algorithmic and scientific advances and considerably accelerate machine-learning model development, validation and transition into biomedical and clinical implementation. TDC is an open-science initiative available at https://tdcommons.ai.","url_abs":"https://arxiv.org/abs/2102.09548v2","url_pdf":"https://arxiv.org/pdf/2102.09548v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"therapeutics-data-commons-machine-learning","repo_url":"https://github.com/mims-harvard/TDC","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"therapeutics-data-commons-machine-learning","repo_url":"https://github.com/yzhao062/yzhao062","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"drug-discovery","task_name":"Drug Discovery"},{"task_slug":"molecular-property-prediction","task_name":"Molecular Property Prediction"},{"task_slug":"tdc-admet-benchmarking-group","task_name":"TDC ADMET Benchmarking Group"},{"task_slug":"therapeutics-data-commons","task_name":"Therapeutics Data Commons"}],"methods":[],"datasets_introduced":[{"slug":"tdcommons","name":"tdcommons","full_name":"Therapeutics Data Commons"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/molecular-property-prediction-on-bbbp-1","task":"Molecular Property Prediction","dataset":"BBBP","model":"AttrMasking","rank_in_archive_order":7,"of":29,"metrics":{"ROC-AUC":"89.2"},"uses_additional_data":false},{"leaderboard":"/sota/molecular-property-prediction-on-bbbp-1","task":"Molecular Property Prediction","dataset":"BBBP","model":"AttentiveFP","rank_in_archive_order":10,"of":29,"metrics":{"ROC-AUC":"85.5"},"uses_additional_data":false},{"leaderboard":"/sota/tdc-admet-benchmarking-group-on-tdcommons","task":"TDC ADMET Benchmarking Group","dataset":"tdcommons","model":"MLP-RDKit2D","rank_in_archive_order":4,"of":12,"metrics":{"TDC.AMES":"0.823","TDC.BBB_Martins":"0.889","TDC.Bioavailability_Ma":"0.672","TDC.CYP2C9_Inhibition_Veith":"0.742","TDC.CYP2C9_Substrate_CarbonMangels":"0.360","TDC.CYP2D6_Inhibition_Veith":"0.616","TDC.CYP2D6_Substrate_CarbonMangels":"0.677","TDC.CYP3A4_Inhibition_Veith":"0.829","TDC.CYP3A4_Substrate_CarbonMangels":"0.639","TDC.Caco2_Wang":"0.393","TDC.Clearance_Hepatocyte_AZ":"0.382","TDC.Clearance_Microsome_AZ":"0.586","TDC.DILI":"0.875","TDC.HIA_Hou":"0.972","TDC.Half_Life_Obach":"0.184","TDC.LD50_Zhu":"0.678","TDC.Lipophilicity_AstraZeneca":"0.574","TDC.PPBR_AZ":"9.994","TDC.Pgp_Broccatelli":"0.918","TDC.Solubility_AqSolDB":"0.827","TDC.VDss_Lombardo":"0.561","TDC.hERG":"0.841"},"uses_additional_data":false},{"leaderboard":"/sota/tdc-admet-benchmarking-group-on-tdcommons","task":"TDC ADMET Benchmarking Group","dataset":"tdcommons","model":"AttentiveFP","rank_in_archive_order":5,"of":12,"metrics":{"TDC.AMES":"0.814","TDC.BBB_Martins":"0.855","TDC.Bioavailability_Ma":"0.632","TDC.CYP2C9_Inhibition_Veith":"0.749","TDC.CYP2C9_Substrate_CarbonMangels":"0.375","TDC.CYP2D6_Inhibition_Veith":"0.646","TDC.CYP2D6_Substrate_CarbonMangels":"0.574","TDC.CYP3A4_Inhibition_Veith":"0.851","TDC.CYP3A4_Substrate_CarbonMangels":"0.576","TDC.Caco2_Wang":"0.401","TDC.Clearance_Hepatocyte_AZ":"0.289","TDC.Clearance_Microsome_AZ":"0.365","TDC.DILI":"0.886","TDC.HIA_Hou":"0.974","TDC.Half_Life_Obach":"0.085","TDC.LD50_Zhu":"0.678","TDC.Lipophilicity_AstraZeneca":"0.572","TDC.PPBR_AZ":"9.373","TDC.Pgp_Broccatelli":"0.892","TDC.Solubility_AqSolDB":"0.776","TDC.VDss_Lombardo":"0.241","TDC.hERG":"0.825"},"uses_additional_data":false},{"leaderboard":"/sota/tdc-admet-benchmarking-group-on-tdcommons","task":"TDC ADMET Benchmarking Group","dataset":"tdcommons","model":"AttrMasking","rank_in_archive_order":6,"of":12,"metrics":{"TDC.AMES":"0.842","TDC.BBB_Martins":"0.892","TDC.Bioavailability_Ma":"0.577","TDC.CYP2C9_Inhibition_Veith":"0.829","TDC.CYP2C9_Substrate_CarbonMangels":"0.381","TDC.CYP2D6_Inhibition_Veith":"0.721","TDC.CYP2D6_Substrate_CarbonMangels":"0.704","TDC.CYP3A4_Inhibition_Veith":"0.902","TDC.CYP3A4_Substrate_CarbonMangels":"0.582","TDC.Caco2_Wang":"0.546","TDC.Clearance_Hepatocyte_AZ":"0.413","TDC.Clearance_Microsome_AZ":"0.585","TDC.DILI":"0.919","TDC.HIA_Hou":"0.978","TDC.Half_Life_Obach":"0.151","TDC.LD50_Zhu":"0.685","TDC.Lipophilicity_AstraZeneca":"0.547","TDC.PPBR_AZ":"10.075","TDC.Pgp_Broccatelli":"0.929","TDC.Solubility_AqSolDB":"1.026","TDC.VDss_Lombardo":"0.559","TDC.hERG":"0.778"},"uses_additional_data":false},{"leaderboard":"/sota/tdc-admet-benchmarking-group-on-tdcommons","task":"TDC ADMET Benchmarking Group","dataset":"tdcommons","model":"GCN","rank_in_archive_order":7,"of":12,"metrics":{"TDC.AMES":"0.818","TDC.BBB_Martins":"0.842","TDC.Bioavailability_Ma":"0.566","TDC.CYP2C9_Inhibition_Veith":"0.735","TDC.CYP2C9_Substrate_CarbonMangels":"0.344","TDC.CYP2D6_Inhibition_Veith":"0.616","TDC.CYP2D6_Substrate_CarbonMangels":"0.617","TDC.CYP3A4_Inhibition_Veith":"0.840","TDC.CYP3A4_Substrate_CarbonMangels":"0.590","TDC.Caco2_Wang":"0.599","TDC.Clearance_Hepatocyte_AZ":"0.366","TDC.Clearance_Microsome_AZ":"0.532","TDC.DILI":"0.859","TDC.HIA_Hou":"0.936","TDC.Half_Life_Obach":"0.239","TDC.LD50_Zhu":"0.649","TDC.Lipophilicity_AstraZeneca":"0.541","TDC.PPBR_AZ":"10.194","TDC.Pgp_Broccatelli":"0.895","TDC.Solubility_AqSolDB":"0.907","TDC.VDss_Lombardo":"0.457","TDC.hERG":"0.738"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2102.09548","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2102.09548"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yzhao062/yzhao062","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mims-harvard/TDC","reach":null}],"summary":{"ran_honours":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"758ec25d48fcc6c8","entry":"latest_ckpt","repo":"mims-harvard/TDC","repo_kind":"official","path":"examples/generation/docking_generation/moldqn/chemgraph/denovo.py","file_url":"https://github.com/mims-harvard/TDC/blob/HEAD/examples/generation/docking_generation/moldqn/chemgraph/denovo.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"758ec25d48fcc6c8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}