{"url":"/dataset/conll-2003","name":"CoNLL 2003","full_name":null,"description_markdown":"**CoNLL-2003** is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.\r\nThe data consists of eight files covering two languages: English and German.\r\nFor each of the languages there is a training file, a development file, a test file and a large file with unannotated data.\r\n\r\nThe English data was taken from the Reuters Corpus. This corpus consists of Reuters news stories between August 1996 and August 1997.\r\nFor the training and development set, ten days worth of data were taken from the files representing the end of August 1996.\r\nFor the test set, the texts were from December 1996. The preprocessed raw data covers the month of September 1996.\r\n\r\nThe text for the German data was taken from the ECI Multilingual Text Corpus. This corpus consists of texts in many languages. The portion of data that\r\nwas used for this task, was extracted from the German newspaper Frankfurter Rundshau. All three of the training, development and test sets were taken\r\nfrom articles written in one week at the end of August 1992.\r\nThe raw data were taken from the months of September to December 1992.\r\n\r\n\r\n| English      data | Articles | Sentences | Tokens  | LOC  | MISC | ORG  | PER  |\r\n|-------------------|----------|-----------|---------|------|------|------|------|\r\n| Training     set  | 946      | 14,987    | 203,621 | 7140 | 3438 | 6321 | 6600 |\r\n| Development  set  | 216      | 3,466     | 51,362  | 1837 | 922  | 1341 | 1842 |\r\n| Test         set  | 231      | 3,684     | 46,435  | 1668 | 702  | 1661 | 1617 |\r\n\r\nNumber of articles, sentences, tokens and entities (locations, miscellaneous, organizations, and persons) in English data files.\r\n\r\n\r\n\r\n| German       data | Articles | Sentences | Tokens  | LOC  | MISC | ORG  | PER  |\r\n|-------------------|----------|-----------|---------|------|------|------|------|\r\n| Training     set  | 553      | 12,705    | 206,931 | 4363 | 2288 | 2427 | 2773 |\r\n| Development  set  | 201      | 3,068     | 51,444  | 1181 | 1010 | 1241 | 1401 |\r\n| Test         set  | 155      | 3,160     | 51,943  | 1035 | 670  | 773  | 1195 |\r\n\r\nNumber of articles, sentences, tokens and entities (locations, miscellaneous, organizations, and persons) in German data files.","description_withheld":null,"homepage":"https://www.clips.uantwerpen.be/conll2003/ner/","introduced_date":"2003-06-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/introduction-to-the-conll-2003-shared-task","title":"Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition","first_author":"Erik F. Tjong Kim Sang","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Named Entity Recognition (NER)","url":"/task/named-entity-recognition-ner","datasets_with_task":"/datasets/task/named-entity-recognition-ner"},{"name":"Information Retrieval","url":"/task/information-retrieval","datasets_with_task":"/datasets/task/information-retrieval"},{"name":"Cross-Lingual NER","url":"/task/cross-lingual-ner","datasets_with_task":"/datasets/task/cross-lingual-ner"},{"name":"UIE","url":"/task/uie","datasets_with_task":"/datasets/task/uie"},{"name":"Token Classification","url":"/task/token-classification","datasets_with_task":"/datasets/task/token-classification"},{"name":"Named Entity Recognition","url":"/task/named-entity-recognition-1","datasets_with_task":"/datasets/task/named-entity-recognition-1"},{"name":"Semantic Similarity","url":"/task/semantic-similarity","datasets_with_task":"/datasets/task/semantic-similarity"},{"name":"Weakly-Supervised Named Entity Recognition","url":"/task/weakly-supervised-named-entity-recognition","datasets_with_task":"/datasets/task/weakly-supervised-named-entity-recognition"},{"name":"Chunking","url":"/task/chunking","datasets_with_task":"/datasets/task/chunking"},{"name":"NER","url":"/task/cg","datasets_with_task":"/datasets/task/cg"},{"name":"Low Resource Named Entity Recognition","url":"/task/low-resource-named-entity-recognition","datasets_with_task":"/datasets/task/low-resource-named-entity-recognition"},{"name":"POS","url":"/task/pos","datasets_with_task":"/datasets/task/pos"},{"name":"FG-1-PG-1","url":"/task/fg-1-pg-1","datasets_with_task":"/datasets/task/fg-1-pg-1"},{"name":"Sparse Information Retrieval","url":"/task/sparse-information-retrieval","datasets_with_task":"/datasets/task/sparse-information-retrieval"},{"name":"Col BERTTriplet","url":"/task/col-berttriplet","datasets_with_task":"/datasets/task/col-berttriplet"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"German","url":"/datasets/language/german"}],"variants":["Unknown","ConLL 2003","conll2003","CoNLL 2003 NER dev","CoNLL 2003 (German) Revised","CONLL 2003 German","CoNLL 2003 (German)","CoNLL 2003 (English)","CoNLL03","CoNLL 2003"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/conll2003","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/tomaarsen/conll2003","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/mohammedriza-rahman/conll2003","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/eriktks/conll2003","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/allenai/allennlp","url":"http://docs.allennlp.org/main/api/data/dataset_readers/conll2003/","frameworks":["pytorch"]}],"num_papers_in_archive":755,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/cross-lingual-ner-on-conll-2003","task":"Cross-Lingual NER","dataset_variant":"CoNLL 2003","rows":4,"metrics":["Spanish","German","Dutch"],"first_row_in_archive_order":{"model":"XLM-RoBERTa-large","paper":"/paper/model-and-data-transfer-for-cross-lingual","metrics":{"Dutch":"82.3","German":"74.5","Spanish":"79.5"},"code_links":[{"title":"ikergarcia1996/Easy-Translate","url":"https://github.com/ikergarcia1996/Easy-Translate"},{"title":"ikergarcia1996/easy-label-projection","url":"https://github.com/ikergarcia1996/easy-label-projection"},{"title":"ikergarcia1996/Iker-Garcia-Ferrero","url":"https://github.com/ikergarcia1996/Iker-Garcia-Ferrero"},{"title":"ikergarcia1996/annotation-projection-app","url":"https://github.com/ikergarcia1996/annotation-projection-app"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/chunking-on-conll-2003","task":"Chunking","dataset_variant":"CoNLL 2003","rows":1,"metrics":["AUC","Accuracy","F1","Precision","Recall"],"first_row_in_archive_order":{"model":"Def2Vec","paper":"/paper/def2vec-extensible-word-embeddings-from","metrics":{"AUC":"93.07","Accuracy":"77.69","F1":"81.45","Precision":"86.56","Recall":"77.69"},"code_links":[{"title":"IreneMorazzoni/def_2_vec_irene","url":"https://github.com/IreneMorazzoni/def_2_vec_irene"},{"title":"vincenzo-scotti/def_2_vec","url":"https://github.com/vincenzo-scotti/def_2_vec"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/ner-on-conll-2003-1","task":"NER","dataset_variant":"CoNLL 2003","rows":1,"metrics":["AUC","Accuracy","F1","Precision","Recall"],"first_row_in_archive_order":{"model":"Def2Vec","paper":"/paper/def2vec-extensible-word-embeddings-from","metrics":{"AUC":"96.28","Accuracy":"71.98","F1":"83.09","Precision":"99.28","Recall":"71.98"},"code_links":[{"title":"IreneMorazzoni/def_2_vec_irene","url":"https://github.com/IreneMorazzoni/def_2_vec_irene"},{"title":"vincenzo-scotti/def_2_vec","url":"https://github.com/vincenzo-scotti/def_2_vec"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/uie-on-conll-2003","task":"UIE","dataset_variant":"CoNLL 2003","rows":1,"metrics":["F1 score"],"first_row_in_archive_order":{"model":"KnowCoder-7b-IE","paper":"/paper/knowcoder-coding-structured-knowledge-into","metrics":{"F1 score":"95.1"},"code_links":[{"title":"ICT-GoKnow/KnowCoder","url":"https://github.com/ICT-GoKnow/KnowCoder"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/information-retrieval-on-unknown","task":"Information Retrieval","dataset_variant":"Unknown","rows":0,"metrics":["Cosine Accuracy@1","Cosine Accuracy@10","Cosine Accuracy@3","Cosine Accuracy@5","Cosine Map@100","Cosine Mrr@10","Cosine Ndcg@10","Cosine Precision@1","Cosine Precision@10","Cosine Precision@3","Cosine Precision@5","Cosine Recall@1","Cosine Recall@10","Cosine Recall@3","Cosine Recall@5"],"first_row_in_archive_order":null,"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/semantic-similarity-on-unknown","task":"Semantic Similarity","dataset_variant":"Unknown","rows":0,"metrics":["Pearson Cosine","Spearman Cosine"],"first_row_in_archive_order":null,"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/knowcoder-coding-structured-knowledge-into","title":"KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction","date":"2024-03-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/def2vec-extensible-word-embeddings-from","title":"Def2Vec: Extensible Word Embeddings from Dictionary Definitions","date":"2023-12-16","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/model-and-data-transfer-for-cross-lingual","title":"Model and Data Transfer for Cross-Lingual Sequence Labelling in Zero-Resource Settings","date":"2022-10-23","rows_on_this_dataset":1,"code_links":4,"syntology":null},{"paper":"/paper/constrained-labeled-data-generation-for-low","title":"Constrained Labeled Data Generation for Low-Resource Named Entity Recognition","date":"2021-08-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/cross-lingual-named-entity-recognition-using","title":"Cross-Lingual Named Entity Recognition Using Parallel Corpus: A New Approach Using XLM-RoBERTa Alignment","date":"2021-01-26","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/entity-projection-via-machine-translation-for","title":"Entity Projection via Machine Translation for Cross-Lingual NER","date":"2019-08-31","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}