Home › Datasets › task › Constituency Parsing

Constituency Parsing datasets

archive 2025-07-28

6 datasets carry the task tag "Constituency Parsing" (the task itself: Constituency Parsing), ordered by the archive's paper count. Page 1 of 1: 6 shown of 6. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Constituency Parsing datasets 1–6 of 6

The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling.
1,006 papers · 10 benchmarks
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
Taiga Corpus (An open-source corpus for machine learning.)
Taiga is a corpus, where text sources and their meta-information are collected according to popular ML tasks.
5 papers · 0 benchmarks
PCC (Potsdam Commentary Corpus)
The Potsdam Commentary Corpus (PCC) is a corpus of 220 German newspaper commentaries (2.900 sentences, 44.000 tokens) taken from the online issues of the Märkische Allgemeine Zeitung (MAZ subcorpus) and Tagesspiegel (ProCon subcorpus) and…
2 papers · 0 benchmarks
The Alexa Point of View dataset is point of view conversion dataset, a parallel corpus of messages spoken to a virtual assistant and the converted messages for delivery.
1 paper · 1 benchmark
MASC (Manually Annotated Sub-Corpus)
The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.