{"url":"/dataset/tox21-1","name":"Tox21","full_name":null,"description_markdown":"The **Tox21** data set comprises 12,060 training samples and 647 test samples that represent chemical compounds. There are 801 \"dense features\" that represent chemical descriptors, such as molecular weight, solubility or surface area, and 272,776 \"sparse features\" that represent chemical substructures (ECFP10, DFS6, DFS8; stored in Matrix Market Format ). Machine learning methods can either use sparse or dense data or combine them. For each sample there are 12 binary labels that represent the outcome (active/inactive) of 12 different toxicological experiments. Note that the label matrix contains many missing values (NAs). The original data source and Tox21 challenge site is https://tripod.nih.gov/tox21/challenge/.\r\n\r\nSource: [Tox21 Machine Learning Data Set](http://bioinf.jku.at/research/DeepTox/tox21.html)\r\nImage Source: [https://www.frontiersin.org/articles/10.3389/fenvs.2015.00080/full](https://www.frontiersin.org/articles/10.3389/fenvs.2015.00080/full)","description_withheld":null,"homepage":"http://bioinf.jku.at/research/DeepTox/tox21.html","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[{"name":"Graph Classification","url":"/task/graph-classification","datasets_with_task":"/datasets/task/graph-classification"},{"name":"Drug Discovery","url":"/task/drug-discovery","datasets_with_task":"/datasets/task/drug-discovery"},{"name":"Graph Regression","url":"/task/graph-regression","datasets_with_task":"/datasets/task/graph-regression"},{"name":"Molecular Property Prediction","url":"/task/molecular-property-prediction","datasets_with_task":"/datasets/task/molecular-property-prediction"},{"name":"Molecular Property Prediction (1-shot))","url":"/task/molecular-property-prediction-1-shot","datasets_with_task":"/datasets/task/molecular-property-prediction-1-shot"}],"languages":[],"variants":["Tox21","Tox21 "],"data_loaders":[],"num_papers_in_archive":30,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/molecular-property-prediction-on-tox21-1","task":"Molecular Property Prediction","dataset_variant":"Tox21","rows":20,"metrics":["ROC-AUC"],"first_row_in_archive_order":{"model":"Deep-CBN","paper":"/paper/integrating-convolutional-layers-and-biformer","metrics":{"ROC-AUC":"92.4"},"code_links":[{"title":"akianfar/Deep-CBN","url":"https://github.com/akianfar/Deep-CBN"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/drug-discovery-on-tox21","task":"Drug Discovery","dataset_variant":"Tox21","rows":11,"metrics":["AUC"],"first_row_in_archive_order":{"model":"elEmBERT-V1","paper":"/paper/structure-to-property-chemical-element","metrics":{"AUC":"0.961"},"code_links":[{"title":"dmamur/elembert","url":"https://github.com/dmamur/elembert"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/graph-classification-on-tox21","task":"Graph Classification","dataset_variant":"Tox21","rows":3,"metrics":["ROC-AUC"],"first_row_in_archive_order":{"model":"GMT","paper":"/paper/accurate-learning-of-graph-representations-1","metrics":{"ROC-AUC":"77.3"},"code_links":[{"title":"JinheonBaek/GMT","url":"https://github.com/JinheonBaek/GMT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/graph-regression-on-tox21","task":"Graph Regression","dataset_variant":"Tox21","rows":3,"metrics":["AUC@80%Train"],"first_row_in_archive_order":{"model":"CensNet","paper":"/paper/censnet-convolution-with-edge-node-switching","metrics":{"AUC@80%Train":"0.78"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/integrating-convolutional-layers-and-biformer","title":"Integrating convolutional layers and biformer network with forward-forward and backpropagation training","date":"2025-02-28","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/pre-training-graph-neural-networks-on","title":"Pre-training Graph Neural Networks on Molecules by Using Subgraph-Conditioned Graph Information Bottleneck","date":"2025-02-20","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/perforated-backpropagation-a-neuroscience","title":"Perforated Backpropagation: A Neuroscience Inspired Extension to Artificial Neural Networks","date":"2025-01-29","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/fine-tuning-graph-neural-networks-by","title":"Fine-tuning Graph Neural Networks by Preserving Graph Generative Patterns","date":"2023-12-21","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":13,"samples_ran":13,"samples_unverified":0,"pointer_only_for_licence":13,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/bioact-het-a-heterogeneous-siamese-neural","title":"BioAct-Het: A Heterogeneous Siamese Neural Network for Bioactivity Prediction Using Novel Bioactivity Representatio","date":"2023-10-15","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/structure-to-property-chemical-element","title":"Structure to Property: Chemical Element Embeddings and a Deep Learning Approach for Accurate Prediction of Chemical Properties","date":"2023-09-17","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/git-mol-a-multi-modal-large-language-model","title":"GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text","date":"2023-08-14","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/molxpt-wrapping-molecules-with-text-for","title":"MolXPT: Wrapping Molecules with Text for Generative Pre-training","date":"2023-05-18","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/galactica-a-large-language-model-for-science-1","title":"Galactica: A Large Language Model for Science","date":"2022-11-16","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/uni-mol-a-universal-3d-molecular","title":"Uni-Mol: A Universal 3D Molecular Representation Learning Framework","date":"2022-09-08","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/fine-tuning-graph-neural-networks-via-graph","title":"Fine-Tuning Graph Neural Networks via Graph Topology induced Optimal Transport","date":"2022-03-20","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":2,"samples_ran":0,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/chemrl-gem-geometry-enhanced-molecular","title":"ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction","date":"2021-06-11","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/accurate-learning-of-graph-representations-1","title":"Accurate Learning of Graph Representations with Graph Multiset Pooling","date":"2021-02-23","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":5,"samples_ran":4,"samples_unverified":1,"pointer_only_for_licence":5,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/trimnet-learning-molecular-representation","title":"TrimNet: learning molecular representation from triplet messages for biomedicine","date":"2020-11-04","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/grover-self-supervised-message-passing","title":"Self-Supervised Graph Transformer on Large-Scale Molecular Data","date":"2020-06-18","rows_on_this_dataset":2,"code_links":3,"syntology":null},{"paper":"/paper/autogluon-tabular-robust-and-accurate-automl","title":"AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data","date":"2020-03-13","rows_on_this_dataset":1,"code_links":7,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":11,"samples_ran":4,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/censnet-convolution-with-edge-node-switching","title":"CensNet: Convolution with Edge-Node Switching in Graph Neural Networks","date":"2019-08-10","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/all-smiles-vae","title":"All SMILES Variational Autoencoder","date":"2019-05-30","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/pre-training-graph-neural-networks","title":"Strategies for Pre-training Graph Neural Networks","date":"2019-05-29","rows_on_this_dataset":2,"code_links":11,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":14,"samples_ran":10,"samples_unverified":4,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/are-learned-molecular-representations-ready","title":"Analyzing Learned Molecular Representations for Property Prediction","date":"2019-04-02","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":15,"samples_ran":10,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/relational-pooling-for-graph-representations","title":"Relational Pooling for Graph Representations","date":"2019-03-06","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/n-gram-graph-a-novel-molecule-representation","title":"N-Gram Graph: Simple Unsupervised Representation for Graphs, with Applications to Molecules","date":"2018-06-24","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":14,"samples_ran":0,"samples_unverified":14,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/toxicblend-virtual-screening-of-toxic","title":"ToxicBlend: Virtual Screening of Toxic Compounds with Ensemble Predictors","date":"2018-06-12","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/learning-graph-level-representation-for-drug","title":"Learning Graph-Level Representation for Drug Discovery","date":"2017-09-12","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/self-normalizing-neural-networks","title":"Self-Normalizing Neural Networks","date":"2017-06-08","rows_on_this_dataset":1,"code_links":13,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/low-data-drug-discovery-with-one-shot","title":"Low Data Drug Discovery with One-shot Learning","date":"2016-11-10","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/convolutional-networks-on-graphs-for-learning","title":"Convolutional Networks on Graphs for Learning Molecular Fingerprints","date":"2015-09-30","rows_on_this_dataset":1,"code_links":8,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":26,"samples_ran":0,"samples_unverified":26,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":12,"samples_harvested":107,"samples_ran":47,"samples_unverified":60,"pointer_only_for_licence":24,"papers_with_no_sample_that_ran":3,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}