Home › Datasets › task › Key Information Extraction

Key Information Extraction datasets

archive 2025-07-28

12 datasets carry the task tag "Key Information Extraction" (the task itself: Key Information Extraction), ordered by the archive's paper count. Page 1 of 1: 12 shown of 12. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Key Information Extraction datasets 1–12 of 12

Consists of a dataset with 1000 whole scanned receipt images and annotations for the competition on scanned receipts OCR and key information extraction (SROIE).
105 papers · 2 benchmarks
CORD (Consolidated Receipt Dataset for Post-OCR Parsing)
OCR is inevitably linked to NLP since its final output is in text.
100 papers · 1 benchmark
EPHOIE (phtnsantader@gmail.com)
EPHOIE is a fully-annotated dataset which is the first Chinese benchmark for both text spotting and visual information extraction.
21 papers · 2 benchmarks
Kleister NDA is a dataset for Key Information Extraction (KIE).
17 papers · 1 benchmark
SIBR (SIBR Dataset for VIE in the Wild)
SIBR是面向自然场景视觉信息抽取的数据集。 1)SIBR总的有1000张图片,400张测试,600张训练,包括中文、英文两种语言。 2)包含images.zip、label.zip、train.txt、test.txt四个文件,images.zip、label.zip中包含所有图片和标签,通过train.txt和test.txt区分训练和测试。…
14 papers · 1 benchmark
DocILE is a large dataset of business documents for the tasks of Key Information Localization and Extraction and Line Item Recognition.
11 papers · 0 benchmarks
Information Extraction from Tables (Extraction materials compositions from tables of materials science research papers)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
The paper used 500 scanned Electronic Theses and Dissertation cover pages (i.e., front pages).
2 papers · 1 benchmark
SIMARA (SIMARA: a database for key-value information extraction from full-page handwritten documents)
Description We propose a new database for information extraction from historical handwritten documents.
2 papers · 2 benchmarks
POIE (Products for OCR and Information Extraction)
Products for OCR and Information Extraction (POIE) dataset derives from camera images of various products in the real world.
1 paper · 0 benchmarks
SOMD (SOftware Mention Detection)
The dataset contains the training and test data for the SOftware Mention Detection challenge.
1 paper · 0 benchmarks
ARF (Artificial Relationships in Fiction)
Artificial Relationships in Fiction Dataset Description Artificial Relationships in Fiction (ARF) is a synthetically annotated dataset for Relation Extraction (RE) in fiction, created from a curated selection of literary texts sourced from…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.