Datasets › CoNLL 2003
CoNLL 2003
CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition. The data consists of eight files covering two languages: English and German. For each of the languages there is a training file, a development file, a test file and a large file with unannotated data.
The English data was taken from the Reuters Corpus. This corpus consists of Reuters news stories between August 1996 and August 1997. For the training and development set, ten days worth of data were taken from the files representing the end of August 1996. For the test set, the texts were from December 1996. The preprocessed raw data covers the month of September 1996.
The text for the German data was taken from the ECI Multilingual Text Corpus. This corpus consists of texts in many languages. The portion of data that was used for this task, was extracted from the German newspaper Frankfurter Rundshau. All three of the training, development and test sets were taken from articles written in one week at the end of August 1992. The raw data were taken from the months of September to December 1992.
| English data | Articles | Sentences | Tokens | LOC | MISC | ORG | PER |
|---|---|---|---|---|---|---|---|
| Training set | 946 | 14,987 | 203,621 | 7140 | 3438 | 6321 | 6600 |
| Development set | 216 | 3,466 | 51,362 | 1837 | 922 | 1341 | 1842 |
| Test set | 231 | 3,684 | 46,435 | 1668 | 702 | 1661 | 1617 |
Number of articles, sentences, tokens and entities (locations, miscellaneous, organizations, and persons) in English data files.
| German data | Articles | Sentences | Tokens | LOC | MISC | ORG | PER |
|---|---|---|---|---|---|---|---|
| Training set | 553 | 12,705 | 206,931 | 4363 | 2288 | 2427 | 2773 |
| Development set | 201 | 3,068 | 51,444 | 1181 | 1010 | 1241 | 1401 |
| Test set | 155 | 3,160 | 51,943 | 1035 | 670 | 773 | 1195 |
Number of articles, sentences, tokens and entities (locations, miscellaneous, organizations, and persons) in German data files.
Benchmarks archive 2025-07-28
All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Cross-Lingual NER | CoNLL 2003 | XLM-RoBERTa-large Spanish 79.5 | Model and Data Transfer for Cross-Lingual Sequence... | ikergarcia1996/Easy-Translate +3 | 4 | Compare |
| Chunking | CoNLL 2003 | Def2Vec AUC 93.07 | Def2Vec: Extensible Word Embeddings from Dictionary Definitions | IreneMorazzoni/def_2_vec_irene +1 | 1 | Compare |
| NER | CoNLL 2003 | Def2Vec AUC 96.28 | Def2Vec: Extensible Word Embeddings from Dictionary Definitions | IreneMorazzoni/def_2_vec_irene +1 | 1 | Compare |
| UIE | CoNLL 2003 | KnowCoder-7b-IE F1 score 95.1 | KnowCoder: Coding Structured Knowledge into LLMs for... | ICT-GoKnow/KnowCoder | 1 | Compare |
| Information Retrieval | Unknown | no rows | — | — | 0 | Compare |
| Semantic Similarity | Unknown | no rows | — | — | 0 | Compare |
Papers archive 2025-07-28
6 shown of 6 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 755. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction | 1 | 1 | 12 Mar 2024 | not harvested |
| Def2Vec: Extensible Word Embeddings from Dictionary Definitions | 2 | 2 | 16 Dec 2023 | not harvested |
| Model and Data Transfer for Cross-Lingual Sequence Labelling in Zero-Resource Settings | 4 | 1 | 23 Oct 2022 | not harvested |
| Constrained Labeled Data Generation for Low-Resource Named Entity Recognition | 0 | 1 | 1 Aug 2021 | not harvested |
| Cross-Lingual Named Entity Recognition Using Parallel Corpus: A New Approach Using XLM-RoBERTa Alignment | 0 | 1 | 26 Jan 2021 | not harvested |
| Entity Projection via Machine Translation for Cross-Lingual NER | 1 | 1 | 31 Aug 2019 | not harvested |
Dataset loaders archive 2025-07-28
5 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Unknown
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Unknown
- ConLL 2003
- conll2003
- CoNLL 2003 NER dev
- CoNLL 2003 (German) Revised
- CONLL 2003 German
- CoNLL 2003 (German)
- CoNLL 2003 (English)
- CoNLL03
- CoNLL 2003
10 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections