{"url":"/dataset/scigen","name":"SciGen","full_name":null,"description_markdown":"**SciGen** is a challenge dataset for the task of reasoning-aware data-to-text generation consisting of tables from scientific articles and their corresponding descriptions.  The unique properties of SciGen are that (1) tables mostly contain numerical values, and (2) the corresponding descriptions require arithmetic reasoning. SciGen is therefore the first dataset that assesses the arithmetic reasoning capabilities of generation models on complex input structures, i.e., tables from scientific articles. SciGen opens new avenues for future research in reasoning-aware text generation and evaluation.\r\n\r\nThe dataset consists of 1.3K pairs of tables with their descriptions, with an average of 53 cells in each table.","description_withheld":null,"homepage":"https://github.com/UKPLab/SciGen/tree/main/dat","introduced_date":"2021-04-16","introduced_date_note":null,"introduced_by":{"paper":"/paper/learning-to-reason-for-text-generation-from","title":"Learning to Reason for Text Generation from Scientific Tables","first_author":"Nafise Sadat Moosavi","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Data-to-Text Generation","url":"/task/data-to-text-generation","datasets_with_task":"/datasets/task/data-to-text-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SciGen"],"data_loaders":[{"repo":"https://github.com/UKPLab/SciGen","url":"https://github.com/UKPLab/SciGen","frameworks":[]}],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}