Home › Datasets › task › Dependency Parsing

Dependency Parsing datasets

archive 2025-07-28

15 datasets carry the task tag "Dependency Parsing" (the task itself: Dependency Parsing), ordered by the archive's paper count. Page 1 of 1: 15 shown of 15. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Dependency Parsing datasets 1–15 of 15

The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling.
1,006 papers · 10 benchmarks
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
English Web Treebank is a dataset containing 254,830 word-level tokens and 16,624 sentence-level tokens of webtext in 1174 files annotated for sentence- and word-level tokenization, part-of-speech, and syntactic structure.
42 papers · 0 benchmarks
KLUE (Korean Language Understanding Evaluation)
Korean Language Understanding Evaluation (KLUE) benchmark is a series of datasets to evaluate natural language understanding capability of Korean language models.
21 papers · 1 benchmark
Briefly describe the dataset.
20 papers · 2 benchmarks
Over 4 million frames of motion capture data for 100 different styles of locomotion.
15 papers · 1 benchmark
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
The task builds on the CoNLL-2008 task and extends it to multiple languages.
8 papers · 2 benchmarks
DaNE (Danish Dependency Treebank)
Danish Dependency Treebank (DaNE) is a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme.
6 papers · 5 benchmarks
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
5 papers · 1 benchmark
WiLI-2018 is a benchmark dataset for monolingual written natural language identification.
3 papers · 0 benchmarks
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by UDPipe (http://ufal.mff.cuni.cz/udpipe), together with word embeddings of dimension 100 computed from lowercased…
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.