Home › Datasets › task › Token Classification
Token Classification datasets
archive 2025-07-28
15 datasets carry the task tag "Token Classification" (the task itself: Token Classification), ordered by the archive's paper count. Page 1 of 1: 15 shown of 15. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Token Classification datasets 1–15 of 15
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
Consists of a dataset with 1000 whole scanned receipt images and annotations for the competition on scanned receipts OCR and key information extraction (SROIE).
105 papers · 2 benchmarks
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
XTREME (Cross-Lingual Transfer Evaluation of Multilingual Encoders)
The Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark was introduced to encourage more research on multilingual transfer learning,.
55 papers · 2 benchmarks
The IndoNLU benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems for Bahasa Indonesia.
14 papers · 2 benchmarks
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
LeNER-Br is a dataset for named entity recognition (NER) in Brazilian Legal Text.
3 papers · 2 benchmarks
GeoEDdA (A Gold Standard Dataset for Geo-semantic Annotation of Diderot & d’Alembert’s Encyclopédie)
Dataset Description - Authors: Ludovic Moncla, Katherine McDonough and Denis Vigier in the framework of the GEODE project.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The MIM-GOLD-NER dataset is an Icelandic named entity (NE) corpus.
0 papers · 1 benchmark
NERGRIT involves machine learning based NLP Tools and a corpus used for Indonesian Named Entity Recognition, Statement Extraction, and Sentiment Analysis.
0 papers · 1 benchmark
The Tagalog Universal Dependencies NewsCrawl dataset consists of annotated text extracted from the Leipzig Tagalog Corpus.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.