{"url":"/dataset/tablex","name":"TabLeX","full_name":null,"description_markdown":"**TabLeX** is a large-scale benchmark dataset comprising table images generated from scientific articles. TabLeX consists of two subsets, one for table structure extraction and the other for table content extraction. Each table image is accompanied by its corresponding LATEX source code. To facilitate the development of robust table IE tools, TabLeX contains images in different aspect ratios and in a variety of fonts.","description_withheld":null,"homepage":"https://drive.google.com/drive/folders/1l60XunwuTShRiJnTbmM8eloECNLXxEPp","introduced_date":"2021-05-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/tablex-a-benchmark-dataset-for-structure-and","title":"TabLeX: A Benchmark Dataset for Structure and Content Information Extraction from Scientific Tables","first_author":"Harsh Desai","url":null},"license":{"name":"Attribution 4.0 International","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[],"languages":[],"variants":["TabLeX"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}