{"url":"/dataset/segmentedtables","name":"SegmentedTables","full_name":null,"description_markdown":"The **SegmentedTables** dataset is a collection of almost 2,000 tables extracted from 352 machine learning papers. Each table consists of rich text content, layout and caption. Tables are annotated with types (leaderboard, ablation, irrelevant) and cells of relevant tables are annotated with semantic roles (such as “paper model”, “competing model”, “dataset”, “metric”).\n\nDue to the license of source papers the dataset is published as a metadata, all annotations and open-source pipeline that can be used to extract the tables.\n\nSource: [AxCell: Automatic Extraction of Results from Machine Learning Papers](https://paperswithcode.com/paper/axcell-automatic-extraction-of-results-from)\nImage Source: [AxCell: Automatic Extraction of Results from Machine Learning Papers](https://paperswithcode.com/paper/axcell-automatic-extraction-of-results-from)","description_withheld":null,"homepage":"https://github.com/paperswithcode/axcell","introduced_date":null,"introduced_date_note":null,"introduced_by":{"paper":"/paper/axcell-automatic-extraction-of-results-from","title":"AxCell: Automatic Extraction of Results from Machine Learning Papers","first_author":"Marcin Kardas","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Tables","url":"/datasets/modality/tables"}],"tasks":[{"name":"Scientific Results Extraction","url":"/task/scientific-results-extraction","datasets_with_task":"/datasets/task/scientific-results-extraction"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SegmentedTables"],"data_loaders":[{"repo":"https://github.com/paperswithcode/axcell","url":"https://github.com/paperswithcode/axcell","frameworks":[]}],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}