Home › Datasets › task › Open Information Extraction

Open Information Extraction datasets

archive 2025-07-28

14 datasets carry the task tag "Open Information Extraction" (the task itself: Open Information Extraction), ordered by the archive's paper count. Page 1 of 1: 14 shown of 14. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Open Information Extraction datasets 1–14 of 14

The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling.
1,006 papers · 10 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
QA-SRL was proposed as an open schema for semantic roles, in which the relation between an argument and a predicate is expressed as a natural-language question containing the predicate (“Where was someone educated?”) whose answer is the…
43 papers · 0 benchmarks
CaRB (Crowdsourced automatic open Relation extraction Benchmark)
CaRB [Bhardwaj et al., 2019] is developed by re-annotating the dev and test splits of OIE2016 via crowd-sourcing.
35 papers · 1 benchmark
OIE2016 is the first large-scale OpenIE benchmark.
31 papers · 1 benchmark
The GenericsKB contains 3.4M+ generic sentences about the world, i.e., sentences expressing general truths such as "Dogs bark," and "Trees remove carbon dioxide from the atmosphere." Generics are potentially useful as a knowledge source…
26 papers · 0 benchmarks
LSOIE (Large-Scale dataset for Supervised Open Information Extraction)
LSOIE is a large-scale OpenIE data converted from QA-SRL 2.0 in two domains, i.e., Wikipedia and Science.
8 papers · 2 benchmarks
OPIEC (Open Information Extraction Corpus)
OPIEC is an Open Information Extraction (OIE) corpus, constructed from the entire English Wikipedia.
8 papers · 0 benchmarks
We manually performed the task of Open Information Extraction on 5 short documents, elaborating tentative guidelines for the task, and resulting in a ground truth reference of 347 tuples.
8 papers · 1 benchmark
This dataset is a new knowledge-base (KB) of hasPart relationships, extracted from a large corpus of generic statements.
7 papers · 0 benchmarks
BenchIE: a benchmark and evaluation framework for comprehensive evaluation of OIE systems for English, Chinese and German.
3 papers · 1 benchmark
We manually annotate 800 sentences from 80 documents in two domains (Healthcare and Transportation) to form a DocOIE dataset for evaluation.
1 paper · 2 benchmarks
Dataset Card for SemTabNet This dataset accompanies the following paper: Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar…
1 paper · 1 benchmark
The TupleInf Open IE dataset contains Open IE tuples extracted from 263K sentences that were used by the solver in “Answering Complex Questions Using Open Information Extraction” (referred as Tuple KB, T).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.