{"url":"/dataset/adore","name":"ADORE","full_name":"A benchmark dataset for machine learning in ecotoxicology","description_markdown":"ADORE is a benchmark dataset for machine learning for ecotixicology, covering acute aquatic toxicity in three relevant taxonomic groups (fish, crustaceans, and algae). The core dataset describes ecotoxicological experiments and is expanded with phylogenetic and species-specific data on the species as well as chemical properties and molecular representations. Apart from challenging other researchers to try and achieve the best model performances across the whole dataset, we propose specific relevant challenges on subsets of the data and include datasets and splittings corresponding to each of these challenge as well as in-depth characterization and discussion of train-test splitting approaches.\r\n\r\nThe dataset contains acute toxicity data (lethal concentration 50; LC50 or effective concentration 50; EC50) on 2,408 chemicals in 203 different species of algae, crustaceans, and fish. This encompasses a total of 33K data points, 26K of which are on fish (140 species).\r\n\r\nThe task is to predict the ecotoxicological outcome based on historic ecotoxicity data.\r\n\r\nADORE was originally published in Nature ScientificData:\r\nSchür, Christoph, Lilian Gasser, Fernando Perez-Cruz, Kristin Schirmer, and Marco Baity-Jesi. 2023. “A Benchmark Dataset for Machine Learning in Ecotoxicology.” Scientific Data 10 (1): 718. https://doi.org/10.1038/s41597-023-02612-2.","description_withheld":null,"homepage":"https://doi.org/10.1038/s41597-023-02612-2","introduced_date":"2023-10-20","introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Biology","url":"/datasets/modality/biology"},{"name":"Environment","url":"/datasets/modality/environment"}],"tasks":[{"name":"regression","url":"/task/regression-1","datasets_with_task":"/datasets/task/regression-1"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ADORE"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/regression-on-adore","task":"regression","dataset_variant":"ADORE","rows":2,"metrics":["chemical macro-average RMSE","micro-averaged RMSE"],"first_row_in_archive_order":{"model":"RF-ToxPrint","paper":"/paper/machine-learning-based-prediction-of-fish","metrics":{"chemical macro-average RMSE":"0.845"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/machine-learning-based-prediction-of-fish","title":"Machine learning-based prediction of fish acute mortality: implementation, interpretation, and regulatory relevance","date":"2024-06-03","rows_on_this_dataset":2,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}