Datasets › CoLA

CoLA (Corpus of Linguistic Acceptability)

Introduced by Alex Warstadt et al. in Neural Network Acceptability Judgments1 Jan 2018 archive 2025-07-28

The Corpus of Linguistic Acceptability (CoLA) consists of 10657 sentences from 23 linguistics publications, expertly annotated for acceptability (grammaticality) by their original authors. The public version contains 9594 sentences belonging to training and development sets, and excludes 1063 sentences belonging to a held out test set.

Source: https://nyu-mll.github.io/CoLA/ Image Source: https://arxiv.org/pdf/1805.12471.pdf

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 35 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 710. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Not all layers are equally as important: Every Layer Counts BERT 0 4 3 Nov 2023 not harvested
LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot Fine-Tuning 1 1 29 May 2023 not harvested
Can BERT eat RuCoLA? Topological Data Analysis to Explain 2 2 4 Apr 2023 not harvested
tasksource: A Dataset Harmonization Framework for Streamlined NLP Multi-Task Learning and Evaluation 1 1 14 Jan 2023 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
RuCoLA: Russian Corpus of Linguistic Acceptability 1 1 23 Oct 2022 not harvested
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale 4 1 15 Aug 2022 ran 2 of 5 samples (3 unverified)
Acceptability Judgements via Examining the Topology of Attention Maps 1 5 19 May 2022 not harvested
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language 12 1 7 Feb 2022 ran 0 of 6 samples (6 unverified)
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization 2 1 23 Jun 2021 ran 7 of 10 samples (3 unverified)
FNet: Mixing Tokens with Fourier Transforms 12 1 9 May 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
Entailment as Few-Shot Learner 3 1 29 Apr 2021 ran 1 of 3 samples (2 unverified)
How to Train BERT with an Academic Budget 4 1 15 Apr 2021 not harvested
CLEAR: Contrastive Learning for Sentence Representation 0 1 31 Dec 2020 not harvested
RealFormer: Transformer Likes Residual Attention 5 1 21 Dec 2020 not harvested
Mixing ADAM and SGD: a Combined Optimization Method 1 1 16 Nov 2020 not harvested
A Statistical Framework for Low-bitwidth Training of Deep Neural Networks 2 1 27 Oct 2020 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
SqueezeBERT: What can computer vision teach NLP about efficient neural networks? 6 1 19 Jun 2020 ran 0 of 1 samples (1 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
Synthesizer: Rethinking Self-Attention in Transformer Models 1 1 2 May 2020 ran 1 of 1 samples (0 unverified)
Learning to Encode Position for Transformer with Continuous Dynamical Model 1 1 13 Mar 2020 ran 3 of 6 samples (3 unverified; 6 pointer-only for licence)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 5 23 Oct 2019 ran 2 of 31 samples (29 unverified)
Q8BERT: Quantized 8Bit BERT 5 1 14 Oct 2019 ran 3 of 11 samples (8 unverified; 3 pointer-only for licence)
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter 37 1 2 Oct 2019 ran 19 of 27 samples (8 unverified)
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 48 1 26 Sep 2019 ran 46 of 126 samples (80 unverified; 22 pointer-only for licence)
TinyBERT: Distilling BERT for Natural Language Understanding 10 2 23 Sep 2019 ran 0 of 4 samples (4 unverified; 4 pointer-only for licence)
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT 0 1 12 Sep 2019 not harvested
StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding 0 1 13 Aug 2019 not harvested
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding 3 2 29 Jul 2019 ran 0 of 1 samples (1 unverified; 1 pointer-only for licence)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)

The full list of 35 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CoLA
  • CoLA Dev

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections