{"url":"/dataset/semtabnet","name":"SemTabNet","full_name":null,"description_markdown":"# Dataset Card for SemTabNet\r\n\r\nThis dataset accompanies the following [paper](https://arxiv.org/abs/2406.19102):\r\n\r\n```\r\nTitle: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs\r\nAuthors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar Berrospi Ramis, Shubham Gupta, Michele Dolfi, Peter Staar\r\nVenue: Accepted at the NLP4Climate workshop in the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024) \r\n```\r\n\r\nIn this paper, we propose **STATEMENTS** as a new knowledge model for storing quantiative information in a domain agnotic, uniform structure. The task of converting a raw input (table or text) to Statements is called Statement Extraction (SE). The statement extraction task falls under the category of universal information extraction.\r\n\r\n- **Code Repository:** [SemTabNet repository](https://github.com/DS4SD/SemTabNet)\r\n- **Arxiv Paper:** [Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs](https://arxiv.org/abs/2406.19102)\r\n- **Point of Contact:** [IBM Research DeepSearch Team](https://ds4sd.github.io)\r\n\r\n\r\n### Data Splits\r\n\r\nThere are three tasks supported by this dataset. The data for each three task is split in training, validation, and testing set. Additionally, we also provide the original annotations of the raw tables which are used to construct all other data.\r\n\r\n|Task | Train   | Test | Valid |\r\n| ----- | ------ | ----- | ---- |\r\n| SE Direct | 103455 | 11682 | 5445 |\r\n|SE Indirect 1D | 72580 | 8489 | 3821 |\r\n|SE Indirect 2D |  93153 | 22839 | 4903 |\r\n\r\n### Languages\r\n\r\nThe text in the dataset is in English.\r\n\r\n### Source and Annotations\r\n\r\nThe source of this dataset and the annotation strategy is described in the paper.\r\n\r\n\r\n### Citation Information\r\n\r\nArxiv: [https://arxiv.org/abs/2406.19102](https://arxiv.org/abs/2406.19102)","description_withheld":null,"homepage":"https://huggingface.co/datasets/ds4sd/SemTabNet","introduced_date":"2024-06-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/statements-universal-information-extraction","title":"Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs","first_author":"Lokesh Mishra","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Tabular","url":"/datasets/modality/tabular"},{"name":"Tables","url":"/datasets/modality/tables"}],"tasks":[{"name":"Information Extraction","url":"/task/information-extraction","datasets_with_task":"/datasets/task/information-extraction"},{"name":"Open Information Extraction","url":"/task/open-information-extraction","datasets_with_task":"/datasets/task/open-information-extraction"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SemTabNet"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/information-extraction-on-semtabnet","task":"Information Extraction","dataset_variant":"SemTabNet","rows":1,"metrics":["average Tree Similarity Score"],"first_row_in_archive_order":{"model":"T5","paper":"/paper/statements-universal-information-extraction","metrics":{"average Tree Similarity Score":"81.76"},"code_links":[{"title":"ds4sd/semtabnet","url":"https://github.com/ds4sd/semtabnet"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/statements-universal-information-extraction","title":"Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs","date":"2024-06-27","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}