Browse State-of-the-Art › Constituency Grammar Induction
Constituency Grammar Induction
20 papers with code · 4 benchmarks · 2 datasets archive 2025-07-28
Inducing a constituency-based phrase structure grammar.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| PTB Diagnostic ECG Database (24 rows) | Ensemble (Generative MBR) | Ensemble Distillation for Unsupervised Constituency Parsing | code | Syntology ran 16 of 22 samples · 6 unverified | Compare |
| CTB (1 row) | SemInfo-NPCFG (60NT) | Improving Unsupervised Constituency Parsing via Maximizing... | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| SPMRL French (1 row) | SemInfo-SNPCFG (1024NT) | Improving Unsupervised Constituency Parsing via Maximizing... | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| SPMRL German (1 row) | SemInfo-NPCFG (60NT) | Improving Unsupervised Constituency Parsing via Maximizing... | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| Penn Treebank (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
20 shown of 20 papers with code (22 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Oct 2018 7 repositories listed Syntology ran 3 of 12 samples · 9 unverifiedWhen a larger constituent ends, all of the smaller constituents that are nested within it must also be closed.
-
13 Mar 2024 2 repositories listed Syntology ran 4 of 7 samples · 3 unverifiedA syntactic language model (SLM) incrementally generates a sentence with its syntactic tree in a left-to-right manner.
-
1 May 2022 2 repositories listedRecent research found it beneficial to use large state spaces for HMMs and PCFGs.
-
1 Mar 2022 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe propose to use a top-down parser as a model-based pruning method, which also enables parallel encoding during inference.
-
24 Jun 2019 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We study a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context-free grammar.
-
5 Oct 2024 1 repository listedUnsupervised parsing, also known as grammar induction, aims to infer syntactic structure from raw text.
-
3 Oct 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we introduce a novel objective for training unsupervised parsers: maximizing the information between constituent structures and sentence semantics (SemInfo).
-
23 Jul 2024 1 repository listedNeural parameterization has significantly advanced unsupervised grammar induction.
-
23 Oct 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Scaling dense PCFGs to thousands of nonterminals via a low-rank parameterization of the rule probability tensor has been shown to be beneficial for unsupervised parsing.
-
3 Oct 2023 1 repository listed Syntology ran 16 of 22 samples · 6 unverified · 22 pointer-only (licence)We investigate the unsupervised constituency parsing task, which organizes words and phrases of a sentence into a hierarchical structure without using linguistically annotated data.
-
28 Sep 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedMore interestingly, the hierarchical structures induced by ReCAT exhibit strong consistency with human-annotated syntactic trees, indicating good interpretability brought by the CIO layers.
-
5 Oct 2021 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 1 pointer-only (licence)We introduce a method for unsupervised parsing that relies on bootstrapping classifiers to identify if a node dominates a specific span in a sentence.
-
20 Sep 2021 1 repository listedOur experiments find that concreteness is a strong indicator for learning dependency grammars, improving the direct attachment score (DAS) by over 50\% as compared to state-of-the-art models trained on pure text.
-
31 May 2021 1 repository listedNeural lexicalized PCFGs (L-PCFGs) have been shown effective in grammar induction.
-
28 Apr 2021 1 repository listedIn this work, we present a new parameterization form of PCFGs based on tensor decomposition, which has at most quadratic computational complexity in the symbol number and therefore allows us to use a much larger number…
-
25 Sep 2020 1 repository listedIn this work, we study visually grounded grammar induction and learn a constituency parser from both unlabeled text and its visual groundings.
-
1 Jun 2019 1 repository listedWe introduce the deep inside-outside recursive autoencoder (DIORA), a fully-unsupervised method for discovering syntax that simultaneously learns representations for constituents within the induced tree.
-
7 Apr 2019 1 repository listedOn language modeling, unsupervised RNNGs perform as well their supervised counterparts on benchmarks in English and Chinese.
-
28 Aug 2018 1 repository listedIn this work, we propose a novel generative model that jointly learns discrete syntactic structure and continuous word representations in an unsupervised fashion by cascading an invertible neural network with a…
-
2 Nov 2017 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedIn this paper, We propose a novel neural language model, called the Parsing-Reading-Predict Networks (PRPN), that can simultaneously induce the syntactic structure from unannotated sentences and leverage the inferred…
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections