Datasets › SemTabNet

SemTabNet

Introduced by Lokesh Mishra et al. in Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs27 Jun 2024 archive 2025-07-28

Dataset Card for SemTabNet

This dataset accompanies the following paper:

Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs
Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar Berrospi Ramis, Shubham Gupta, Michele Dolfi, Peter Staar
Venue: Accepted at the NLP4Climate workshop in the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024) 

In this paper, we propose STATEMENTS as a new knowledge model for storing quantiative information in a domain agnotic, uniform structure. The task of converting a raw input (table or text) to Statements is called Statement Extraction (SE). The statement extraction task falls under the category of universal information extraction.

Data Splits

There are three tasks supported by this dataset. The data for each three task is split in training, validation, and testing set. Additionally, we also provide the original annotations of the raw tables which are used to construct all other data.

Task Train Test Valid
SE Direct 103455 11682 5445
SE Indirect 1D 72580 8489 3821
SE Indirect 2D 93153 22839 4903
Languages

The text in the dataset is in English.

Source and Annotations

The source of this dataset and the annotation strategy is described in the paper.

Citation Information

Arxiv: https://arxiv.org/abs/2406.19102

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Information Extraction SemTabNet T5 average Tree Similarity Score 81.76 Statements: Universal Information Extraction from Tables... ds4sd/semtabnet 1 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs 1 1 27 Jun 2024 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • SemTabNet

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections