Home › Datasets › task › Event Extraction

Event Extraction datasets

archive 2025-07-28

16 datasets carry the task tag "Event Extraction" (the task itself: Event Extraction), ordered by the archive's paper count. Page 1 of 1: 16 shown of 16. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Event Extraction datasets 1–16 of 16

The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
ACE 2005 (ACE 2005 Multilingual Training Corpus)
ACE 2005 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2005 Automatic Content Extraction (ACE) technology evaluation.
65 papers · 8 benchmarks
SCIREX is a document level IE dataset that encompasses multiple IE tasks, including salient entity identification and document level N-ary relation identification from scientific articles.
33 papers · 2 benchmarks
WikiEvents is a document-level event extraction benchmark dataset which includes complete event and coreference annotation.
33 papers · 1 benchmark
Ten years (2008-2018) ChFinAnn documents and human-summarized event knowledge bases to conduct the DS-based event labeling.
19 papers · 1 benchmark
Aims to extract events and their arguments from multimedia documents.
16 papers · 0 benchmarks
Phee is a dataset for pharmacovigilance comprising over 5000 annotated events from medical case reports and biomedical literature.
8 papers · 0 benchmarks
French TimeBank, a corpus for French annotated in ISO-TimeML.
6 papers · 1 benchmark
The EDT dataset is designed for corporate event detection and text-based stock prediction (trading strategy) benchmark.
5 papers · 0 benchmarks
EventNarrative is a knowledge graph-to-text dataset from publicly available open-world knowledge graphs.
3 papers · 1 benchmark
Title2Event is a large-scale sentence-level dataset for benchmarking Open Event Extraction without restricting event types.
3 papers · 0 benchmarks
Catalan TimeBank 1.0 was developed by researchers at Barcelona Media and consists of Catalan texts in the AnCora corpus annotated with temporal and event information according to the TimeML specification language.
2 papers · 1 benchmark
IndiaPoliceEvents is a corpus of 21,391 sentences from 1,257 English-language Times of India articles about events in the state of Gujarat during March 2002.
2 papers · 0 benchmarks
BKEE (BKEE: Pioneering Event Extraction in the Vietnamese Language)
A novel event extraction dataset for Vietnamese.
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
The PEDC is a corpus of 14 episodes of This American Life podcast transcripts that have been annotated for events.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.