{"url":"/dataset/summaries-of-genetic-variation","name":"Summaries of genetic variation","full_name":null,"description_markdown":"The dataset represents data generated from a commonly used model in population genetics. It comprises a matrix of 1,000,000 rows and 9 columns, representing parameters and summaries generated by an infinite-sites coalescent model for genetic variation. The first two columns encode the scaled mutation rate (theta) and scaled recombination rate (rho). The subsequent seven columns are data summaries: number of segregating sites (C1), standard uniform random noise acting as a distractor (C2), pairwise mean number of nucleotidic differences (C3), mean $R^2$ across pairs separated by <10% of the simulated genomic regions (C4), number of distinct haplotypes (C5), frequency of the most common haplotype (C6), number of singleton haplotypes (C7).\r\n\r\n(this text is not original and adapted from https://journal.r-project.org/archive/2015-2/nunes-prangle.pdf).","description_withheld":null,"homepage":"https://github.com/dennisprangle/abctools","introduced_date":"2010-09-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/on-optimal-selection-of-summary-statistics","title":"On Optimal Selection of Summary Statistics for Approximate Bayesian Computation","first_author":"Matthew A Nunes","url":null},"license":{"name":"MIT","url":"https://choosealicense.com/licenses/mit/"},"modalities":[{"name":"Biology","url":"/datasets/modality/biology"},{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Bayesian Inference","url":"/task/bayesian-inference","datasets_with_task":"/datasets/task/bayesian-inference"}],"languages":[],"variants":["Summaries of genetic variation"],"data_loaders":[{"repo":"https://github.com/tillahoffmann/coaloracle","url":"https://github.com/tillahoffmann/coaloracle","frameworks":[]}],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}