Datasets › CoNLL

CoNLL

archive 2025-07-28

The CoNLL dataset is a widely used resource in the field of natural language processing (NLP). The term “CoNLL” stands for Conference on Natural Language Learning. It originates from a series of shared tasks organized at the Conferences of Natural Language Learning.

Benchmarks archive 2025-07-28

All 35 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Named Entity Recognition (NER) CoNLL 2003 (English) ACE + document-context F1 94.6 Automated Concatenation of Embeddings for Structured Prediction Alibaba-NLP/ACE +1 73 Compare
Grammatical Error Correction CoNLL-2014 Shared Task Ensembles of best 7 models + GRECO + GTP-rerank F0.5 72.8 Pillars of Grammatical Error Correction: Comprehensive... grammarly/pillars-of-gec 23 Compare
Entity Disambiguation AIDA-CoNLL confidence-order In-KB Accuracy 95.0 Global Entity Disambiguation with BERT studio-ousia/luke 20 Compare
Coreference Resolution CoNLL 2012 Maverick_mes Avg F1 83.6 Maverick: Efficient and Accurate Coreference Resolution... sapienzanlp/maverick-coref 18 Compare
Entity Linking AIDA-CoNLL SpEL-large (2023) Micro-F1 strong 88.6 SpEL: Structured Prediction for Entity Linking shavarani/spel 17 Compare
Semantic Role Labeling CoNLL 2005 MRC-SRL F1 90.0 An MRC Framework for Semantic Role Labeling shannonai/mrc-srl 15 Compare
Cross-Lingual NER CoNLL Dutch Zero shot mBERT 3 F1 83.35 Towards Lingua Franca Named Entity Recognition with BERT — 10 Compare
Cross-Lingual NER CoNLL German SMTS Multi sim F1 75.33 Single-/Multi-Source Cross-Lingual NER via... microsoft/vert-papers 10 Compare
Cross-Lingual NER CoNLL Spanish XLM-R large F1 79.5 Model and Data Transfer for Cross-Lingual Sequence... ikergarcia1996/Easy-Translate +3 10 Compare
Chunking CoNLL 2000 ACE Exact Span F1 97.3 Automated Concatenation of Embeddings for Structured Prediction Alibaba-NLP/ACE +1 9 Compare
Grammatical Error Detection CoNLL-2014 A1 VERNet F0.5 54.3 Neural Quality Estimation with Multiple Hypotheses for... thunlp/VERNet 8 Compare
Grammatical Error Detection CoNLL-2014 A2 VERNet F0.5 63.1 Neural Quality Estimation with Multiple Hypotheses for... thunlp/VERNet 8 Compare
Semantic Role Labeling (predicted predicates) CoNLL 2012 Fei et al. 2021 (HeSyFu) + RoBERTa F1 88.59 Better Combine Them Together! Integrating Syntactic... scofield7419/hesyfu 7 Compare
Named Entity Recognition (NER) CoNLL 2003 (German) ACE + document-context F1 88.38 Automated Concatenation of Embeddings for Structured Prediction Alibaba-NLP/ACE +1 6 Compare
Named Entity Recognition (NER) CoNLL 2003 (German) Revised FLERT XLM-R F1 92.23 FLERT: Document-Level Features for Named Entity Recognition flairNLP/flair 5 Compare
Semantic Role Labeling (predicted predicates) CoNLL 2005 LISA + ELMo F1 86.90 Linguistically-Informed Self-Attention for Semantic Role Labeling strubell/LISA 5 Compare
Chunking CoNLL 2003 (English) ACE F1 92.5 Automated Concatenation of Embeddings for Structured Prediction Alibaba-NLP/ACE +1 3 Compare
Chunking CoNLL 2003 (German) ACE F1 95.0 Automated Concatenation of Embeddings for Structured Prediction Alibaba-NLP/ACE +1 3 Compare
Entity Linking CoNLL-Aida RELIC + CoNLL-Aida tuning Accuracy 94.9 Learning Cross-Context Entity Representations from Text — 3 Compare
Grammatical Error Correction CoNLL-2014 Shared Task (10 annotations) GRECO (vote+ESC) F0.5 85.21 System Combination via Quality Estimation for... nusnlp/greco 3 Compare
UCCA Parsing CoNLL 2019 Transition-based (+BERT + Efficient Training + Effective Encoding) Full UCCA F1 66.7 HIT-SCIR at MRP 2019: A Unified Pipeline for Meaning... — 3 Compare
Coreference Resolution CoNLL12 DeepStruct multi-task w/ finetune Average F1 73.1 DeepStruct: Pretraining of Language Models for Structure... cgraywang/deepstruct 2 Compare
Predicate Detection CoNLL 2005 LISA F1 98.4 Linguistically-Informed Self-Attention for Semantic Role Labeling strubell/LISA 2 Compare
Semantic Role Labeling CoNLL05 Brown DeepStruct multi-task w/ finetune F1 92.1 DeepStruct: Pretraining of Language Models for Structure... cgraywang/deepstruct 2 Compare
Semantic Role Labeling CoNLL05 WSJ DeepStruct multi-task F1 95.5 DeepStruct: Pretraining of Language Models for Structure... cgraywang/deepstruct 2 Compare
Semantic Role Labeling CoNLL12 DeepStruct multi-task F1 97.2 DeepStruct: Pretraining of Language Models for Structure... cgraywang/deepstruct 2 Compare
Entity Typing AIDA-CoNLL ReFinED Micro-F1 84.0 ReFinED: An Efficient Zero-shot-capable Approach to... amazon-science/ReFinED +2 1 Compare
FG-1-PG-1 conll2003 CFNER F1 (macro) 0.7911 Distilling Causal Effect from Miscellaneous Other-Class... zzz47zzz/CFNER 1 Compare
Low Resource Named Entity Recognition CONLL 2003 German Zero-Resource Transfer From CoNLL-2003 English dataset. F1 score 65.24 Zero-Resource Cross-Lingual Named Entity Recognition ntunlp/Zero-Shot-Cross-Lingual-NER 1 Compare
Low Resource Named Entity Recognition Conll 2003 Spanish Zero-Resource Cross-lingual Transfer From CoNLL-2003 English dataset. F1 score 75.93 Zero-Resource Cross-Lingual Named Entity Recognition ntunlp/Zero-Shot-Cross-Lingual-NER 1 Compare
Low Resource Named Entity Recognition CONLL 2003 Dutch Zero-Resource Transfer From CoNLL-2003 English dataset. F1 score 74.61 Zero-Resource Cross-Lingual Named Entity Recognition ntunlp/Zero-Shot-Cross-Lingual-NER 1 Compare
Named Entity Recognition (NER) CoNLL 2000 SWEM-CRF F1 90.34 Baseline Needs More Love: On Simple Word-Embedding-Based... dinghanshen/SWEM +1 1 Compare
Predicate Detection CoNLL 2012 LISA F1 97.2 Linguistically-Informed Self-Attention for Semantic Role Labeling strubell/LISA 1 Compare
Semantic Role Labeling CoNLL 2012 HeSyFu Avg. F1 88.59 Better Combine Them Together! Integrating Syntactic... scofield7419/hesyfu 1 Compare
Token Classification conll2003 no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 156 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 187. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Efficient and Interpretable Grammatical Error Correction with Mixture of Experts 1 1 30 Oct 2024 not harvested
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization 1 2 10 Sep 2024 not harvested
ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget 2 2 31 Jul 2024 not harvested
Maverick: Efficient and Accurate Coreference Resolution Defying Recent Trends 1 1 31 Jul 2024 ran 5 of 11 samples (6 unverified; 11 pointer-only for licence)
Transformer-based Named Entity Recognition with Combined Data Representation 0 1 25 Jun 2024 not harvested
Pillars of Grammatical Error Correction: Comprehensive Inspection Of Contemporary Approaches In The Era of Large Language Models 1 2 23 Apr 2024 not harvested
Entity Disambiguation via Fusion Entity Decoding 0 1 2 Apr 2024 not harvested
Unsupervised Grammatical Error Correction Rivaling Supervised Methods 1 1 6 Dec 2023 not harvested
System Combination via Quality Estimation for Grammatical Error Correction 1 2 23 Oct 2023 not harvested
SpEL: Structured Prediction for Entity Linking 1 2 23 Oct 2023 not harvested
Improving Seq2Seq Grammatical Error Correction via Decoding Interventions 1 1 23 Oct 2023 not harvested
GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction 1 1 5 Oct 2023 ran 18 of 25 samples (7 unverified)
PromptNER: Prompt Locating and Typing for Named Entity Recognition 1 2 26 May 2023 not harvested
DiffusionNER: Boundary Diffusion for Named Entity Recognition 3 1 22 May 2023 not harvested
Coreference Resolution through a seq2seq Transition-Based System 1 1 22 Nov 2022 not harvested
Autoregressive Structured Prediction with Language Models 1 3 26 Oct 2022 not harvested
Model and Data Transfer for Cross-Lingual Sequence Labelling in Zero-Resource Settings 4 3 23 Oct 2022 not harvested
SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented Parser 1 1 22 Oct 2022 not harvested
Distilling Causal Effect from Miscellaneous Other-Class for Continual Named Entity Recognition 1 1 8 Oct 2022 not harvested
ReFinED: An Efficient Zero-shot-capable Approach to End-to-End Entity Linking 3 3 8 Jul 2022 not harvested
Improving Entity Disambiguation by Reasoning over a Knowledge Base 3 1 8 Jul 2022 not harvested
Frustratingly Easy System Combination for Grammatical Error Correction 1 1 1 Jul 2022 not harvested
Syntax-driven Approach for Semantic Role Labeling 1 1 1 Jun 2022 not harvested
DeepStruct: Pretraining of Language Models for Structure Prediction 1 8 21 May 2022 ran 7 of 13 samples (6 unverified)
Boundary Smoothing for Named Entity Recognition 1 1 26 Apr 2022 not harvested
Parallel Instance Query Network for Named Entity Recognition 1 1 20 Mar 2022 not harvested
Unified Named Entity Recognition as Word-Word Relation Classification 1 1 19 Dec 2021 not harvested
Named entity recognition architecture combining contextual and global features 1 3 15 Dec 2021 not harvested
Focusing on Potential Named Entities During Active Label Acquisition 1 1 6 Nov 2021 not harvested
Named Entity Recognition for Entity Linking: What Works and What’s Next 1 1 1 Nov 2021 not harvested

The full list of 156 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • AIDA CoNLL-YAGO
  • AIDA-CoNLL
  • wikiann-conll2003
  • tner/conll2003
  • ju-bezdek/conll2003-SK-NER
  • conllpp
  • CoNLL 2019
  • CoNLL 2017 Shared Task - Automatically Annotated Raw Texts and Word Embeddings
  • CoNLL-2014 Shared Task: Grammatical Error Correction
  • CoNLL-2014 Shared Task (10 annotations)
  • CoNLL-2014 10 Annotations
  • CoNLL-2012
  • CoNLL 2003 NER dev
  • CoNLL2003 (English)
  • ConLL 2003
  • conll2003
  • CoNLL 2002
  • conll2002
  • CoNLL-2000
  • CoNLL12
  • CoNLL05 WSJ
  • CoNLL05 Brown
  • CoNLL03
  • CoNLL
  • Conll 2003 Spanish
  • CoNLL04
  • CoNLL-Aida
  • CoNLL-2014 Shared Task
  • CoNLL-2014 A2
  • CoNLL-2014 A1
  • CoNLL-2009
  • CoNLL++
  • CoNLL Spanish
  • CoNLL German
  • CoNLL Dutch
  • CoNLL 2012
  • CoNLL 2005
  • CoNLL 2003 (German) Revised
  • CoNLL 2003 (German)
  • CoNLL 2003 (English)
  • CoNLL 2002 (Spanish)
  • CoNLL 2002 (Dutch)
  • CoNLL 2000
  • CONLL 2003 German
  • CONLL 2003 Dutch

45 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections