{"url":"/dataset/wikitableset","name":"WikiTableSet","full_name":"Wikipedia Table Image Dataset","description_markdown":"WikiTableSet is a large publicly available image-based table recognition dataset in three languages built from Wikipedia.\r\nWikiTableSet contains nearly 4 million English table images, 590K Japanese table images, 640k French table images with corresponding HTML representation, and cell bounding boxes.\r\nWe build a Wikipedia table extractor [WTabHTML](https://github.com/phucty/wtabhtml) and use this to extract tables (in HTML code format) from the 2022-03-01 dump of Wikipedia. In this study, we select Wikipedia tables from three representative languages, i.e., English, Japanese, and French; however, the dataset could be extended to around 300 languages with 17M tables using our table extractor. \r\nSecond, we normalize the HTML tables following the PubTabNet format (separating table headers and table data, removing CSS and style tags). Finally, we use Chrome and Selenium to render table images from table HTML codes. \r\nThis dataset provides a standard benchmark for studying table recognition algorithms in different languages or even multilingual table recognition algorithms.\r\nYou can click [here](https://arxiv.org/pdf/2303.07641.pdf) for more details about this dataset.","description_withheld":null,"homepage":"https://github.com/namtuanly/WikiTableSet","introduced_date":"2023-02-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/rethinking-image-based-table-recognition","title":"Rethinking Image-based Table Recognition Using Weakly Supervised Methods","first_author":"Nam Tuan Ly","url":null},"license":{"name":"MIT","url":"https://github.com/namtuanly/WikiTableSet/blob/main/LICENSE"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Table Recognition","url":"/task/table-recognition","datasets_with_task":"/datasets/task/table-recognition"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"French","url":"/datasets/language/french"},{"name":"Japanese","url":"/datasets/language/japanese"}],"variants":["WikiTableSet"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}