Home › Datasets › task › Entity Linking

Entity Linking datasets

archive 2025-07-28

44 datasets carry the task tag "Entity Linking" (the task itself: Entity Linking), ordered by the archive's paper count. Page 1 of 1: 44 shown of 44. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Entity Linking datasets 1–44 of 44

The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
FUNSD (Form Understanding in Noisy Scanned Documents)
Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms.
179 papers · 3 benchmarks
KILT (KILT Benchmark)
KILT (Knowledge Intensive Language Tasks) is a benchmark consisting of 11 datasets representing 5 types of tasks: Fact-checking (FEVER), Entity linking (AIDA CoNLL-YAGO, WNED-WIKI, WNED-CWEB), Slot filling (T-Rex, Zero Shot RE), Open…
117 papers · 11 benchmarks
FIGER (Fine-Grained Entity Recognition)
The FIGER dataset is an entity recognition dataset where entities are labelled using fine-grained system 112 tags, such as person/doctor, art/writtenwork and building/hotel.
96 papers · 2 benchmarks
AIDA CoNLL-YAGO contains assignments of entities to the mentions of named entities annotated for the original CoNLL 2003 entity recognition task.
64 papers · 0 benchmarks
WebQuestionsSP (WebQuestions Semantic Parses Dataset)
The WebQuestionsSP dataset is released as part of our ACL-2016 paper “The Value of Semantic Parse Labeling for Knowledge Base Question Answering” [Yih, Richardson, Meek, Chang & Suh, 2016], in which we evaluated the value of gathering…
61 papers · 3 benchmarks
MedMentions is a new manually annotated resource for the recognition of biomedical concepts.
48 papers · 1 benchmark
Wikipedia abstracts automatically annotated with WikiData entities and relations that are entailed by the text.
46 papers · 2 benchmarks
IPM NEL (Derczynski IPM Named Entity Linking)
This data is for the task of named entity recognition and linking/disambiguation over tweets.
32 papers · 1 benchmark
BioRED is a first-of-its-kind biomedical relation extraction dataset with multiple entity types (e.g.
25 papers · 3 benchmarks
ZESHEL is a zero-shot entity linking dataset, which places more emphasis on understanding the unstructured descriptions of entities to resolve the ambiguity of mentions on four unseen domains.
25 papers · 1 benchmark
Consists of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph.
21 papers · 0 benchmarks
OVEN (Open-domain Visual Entity Recognition)
In this project, we formally present the task of Open-domain Visual Entity recognitioN (OVEN), where a model need to link an image onto a Wikipedia entity with respect to a text query.
19 papers · 1 benchmark
DWIE (Deutsche Welle corpus for Information Extraction)
The 'Deutsche Welle corpus for Information Extraction' (DWIE) is a multi-task dataset that combines four main Information Extraction (IE) annotation sub-tasks: (i) Named Entity Recognition (NER), (ii) Coreference Resolution, (iii) Relation…
18 papers · 5 benchmarks
FreebaseQA is a data set for open-domain QA over the Freebase knowledge graph.
17 papers · 0 benchmarks
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
MEIR (Multimodal Entity Image Repurposing)
MEIR is a substantially challenging dataset over that which has been previously available to support research into image repurposing detection.
9 papers · 0 benchmarks
WiC-TSV (Words-in-Context: Target Sense Verification)
WiC-TSV is a new multi-domain evaluation benchmark for Word Sense Disambiguation.
9 papers · 2 benchmarks
BB (Bacteria Biotope)
The Bacteria Biotope (BB) Task is part of the BioNLP Open Shared Tasks and meets the BioNLP-OST standards of quality, originality and data formats.
8 papers · 0 benchmarks
A large new multilingual dataset for multilingual entity linking.
8 papers · 1 benchmark
DBLP-QuAD (DBLP Question Answering Dataset)
In this work we create a question answering dataset over the DBLP scholarly knowledge graph (KG).
7 papers · 0 benchmarks
RuBQ (Russian Knowledge Base Questions)
The first Russian knowledge base question answering (KBQA) dataset.
7 papers · 0 benchmarks
KnowledgeNet is a benchmark dataset for the task of automatically populating a knowledge base (Wikidata) with facts expressed in natural language text on the web.
6 papers · 0 benchmarks
XL-BEL is a benchmark for cross-lingual biomedical entity linking (XL-BEL).
6 papers · 0 benchmarks
SoMeSci (Software Mentions in Scientific Articles)
Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling.
5 papers · 0 benchmarks
Twitter-MEL is a multimodal entity linking (MEL) dataset built from Twitter.
5 papers · 0 benchmarks
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
WIKIPerson is a high-quality human-annotated visual person linking dataset based on Wikipedia.
4 papers · 0 benchmarks
EC-FUNSD is introduced in [[arXiv:2402.02379]](https://arxiv.org/abs/2402.02379) as a benchmark of semantic entity recognition (SER) and entity linking (EL), designed for the entity-centric robustness evaluation of pre-trained…
3 papers · 2 benchmarks
AIDA/testc is a new challenging test set for entity linking systems containing 131 Reuters news articles published between December 5th and 7th, 2020.
2 papers · 1 benchmark
BC7 NLM-Chem (BioCreative VII NLM-Chem)
Full-text chemical identification and indexing in PubMed articles.
2 papers · 3 benchmarks
FarsBase-KBP contains 22015 sentences, in which the entities and relation types are linked to the FarsBase ontology.
2 papers · 0 benchmarks
Hansel is a human-annotated Chinese entity linking (EL) dataset, focusing on tail entities and emerging entities: - The test set contains Few-shot (FS) and zero-shot (ZS) slices, has 10K examples and uses Wikidata as the corresponding…
2 papers · 0 benchmarks
LatamXIX (19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction)
A novel dataset of 19th-century Latin American press texts, which addresses the lack of specialized corpora for historical and linguistic analysis in this region.
2 papers · 0 benchmarks
MobIE is a German-language dataset which is human-annotated with 20 coarse- and fine-grained entity types and entity linking information for geographically linkable entities.
2 papers · 0 benchmarks
Rare Diseases Mentions in MIMIC-III (Rare disease mention annotations from a sample of MIMIC-III clinical notes)
Data annotation The 1,073 full rare disease mention annotations (from 312 MIMIC-III discharge summaries) are in fullsetRDannMIMICIIIdisch.csv.
2 papers · 1 benchmark
The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.
1 paper · 0 benchmarks
This dataset contains general and named entities annotations on both clean written text and on noisy speech data.
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
Diderot’s Encyclopédie is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era.
1 paper · 0 benchmarks
This dataset links all the entries describing named entities of Petit Larousse illustré, a French dictionary published in 1905, to wikidata identifiers.
1 paper · 0 benchmarks
Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.