Datasets › SemTabNet
SemTabNet
Dataset Card for SemTabNet
This dataset accompanies the following paper:
Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs
Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar Berrospi Ramis, Shubham Gupta, Michele Dolfi, Peter Staar
Venue: Accepted at the NLP4Climate workshop in the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024)
In this paper, we propose STATEMENTS as a new knowledge model for storing quantiative information in a domain agnotic, uniform structure. The task of converting a raw input (table or text) to Statements is called Statement Extraction (SE). The statement extraction task falls under the category of universal information extraction.
- Code Repository: SemTabNet repository
- Arxiv Paper: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs
- Point of Contact: IBM Research DeepSearch Team
Data Splits
There are three tasks supported by this dataset. The data for each three task is split in training, validation, and testing set. Additionally, we also provide the original annotations of the raw tables which are used to construct all other data.
| Task | Train | Test | Valid |
|---|---|---|---|
| SE Direct | 103455 | 11682 | 5445 |
| SE Indirect 1D | 72580 | 8489 | 3821 |
| SE Indirect 2D | 93153 | 22839 | 4903 |
Languages
The text in the dataset is in English.
Source and Annotations
The source of this dataset and the annotation strategy is described in the paper.
Citation Information
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Information Extraction | SemTabNet | T5 average Tree Similarity Score 81.76 | Statements: Universal Information Extraction from Tables... | ds4sd/semtabnet | 1 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs | 1 | 1 | 27 Jun 2024 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- SemTabNet
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections