Home › Datasets › task › Part-Of-Speech Tagging

Part-Of-Speech Tagging datasets

archive 2025-07-28

26 datasets carry the task tag "Part-Of-Speech Tagging" (the task itself: Part-Of-Speech Tagging), ordered by the archive's paper count. Page 1 of 1: 26 shown of 26. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Part-Of-Speech Tagging datasets 1–26 of 26

The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling.
1,006 papers · 10 benchmarks
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
English Web Treebank is a dataset containing 254,830 word-level tokens and 16,624 sentence-level tokens of webtext in 1174 files annotated for sentence- and word-level tokenization, part-of-speech, and syntactic structure.
42 papers · 0 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
LinCE (Linguistic Code-switching Evaluation Dataset)
A centralized benchmark for Linguistic Code-switching Evaluation (LinCE) that combines ten corpora covering four different code-switched language pairs (i.e., Spanish-English, Nepali-English, Hindi-English, and Modern Standard…
21 papers · 0 benchmarks
Briefly describe the dataset.
20 papers · 2 benchmarks
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
DaNE (Danish Dependency Treebank)
Danish Dependency Treebank (DaNE) is a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme.
6 papers · 5 benchmarks
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
CUGE is a Chinese Language Understanding and Generation Evaluation benchmark with the following features: (1) Hierarchical benchmark framework, where datasets are principally selected and organized with a language capability-task-dataset…
4 papers · 0 benchmarks
Finer (Finnish News Corpus for Named Entity Recognition)
Finnish News Corpus for Named Entity Recognition (Finer) is a corpus that consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, product, event,and date).
4 papers · 0 benchmarks
The Szeged Treebank is the largest fully manually annotated treebank of the Hungarian language.
4 papers · 0 benchmarks
ANTILLES (ANTILLES: An Open French Linguistically Enriched Part-of-Speech Corpus)
ANTILLES is a part-of-speech tagging corpus based on UDFrench-GSD which was originally created in 2015 and is based on the universal dependency treebank v2.0.
1 paper · 1 benchmark
The Alexa Point of View dataset is point of view conversion dataset, a parallel corpus of messages spoken to a virtual assistant and the converted messages for delivery.
1 paper · 1 benchmark
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by UDPipe (http://ufal.mff.cuni.cz/udpipe), together with word embeddings of dimension 100 computed from lowercased…
1 paper · 0 benchmarks
Contains 350 tweets with more than 8,000 words including 3,000 unique words written in Egyptian dialect.
1 paper · 0 benchmarks
IgboNLP is a standard machine translation benchmark dataset for Igbo.
1 paper · 0 benchmarks
MASC (Manually Annotated Sub-Corpus)
The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).
1 paper · 0 benchmarks
This dataset is for evaluation of morphosyntactic analyzers.
1 paper · 1 benchmark
Ritter PoS (Ritter Twitter part-of-speech tagging)
PTB-tagged English Tweets
1 paper · 0 benchmarks
Ensemble Tagger Training and Testing Set This data includes two files: The training set used to create the SCANL Ensemble tagger [1] and the "unseen" testing set that includes words from systems that are not available in the training set.
1 paper · 0 benchmarks
Twitter PoS VCB (Twitter part-of-speech vote-constrained-bootstrapping)
The data is about 1.5 million English tweets annotated for part-of-speech using Ritter's extension of the PTB tagset.
1 paper · 0 benchmarks
Mac-Morpho is a corpus of Brazilian Portuguese texts annotated with part-of-speech tags.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.