Browse State-of-the-Art › Coherence Evaluation
Coherence Evaluation
15 papers with code · 2 benchmarks · 1 dataset archive 2025-07-28
Evaluating the overall coherence of text as measured by its readability and flow through ideas.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| GCDC + RST - Accuracy (4 rows) | MTL with Transformer | Transformer Models for Text Coherence Assessment | code | — | Compare |
| GCDC + RST - F1 (3 rows) | RST-Ensemble | Neural RST-based Evaluation of Discourse Coherence | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
15 shown of 15 papers with code (28 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Sep 2021 2 repositories listedCoherence is an important aspect of text quality and is crucial for ensuring its readability.
-
16 Jul 2024 1 repository listedMotivated by the need for lightweight, open source, and multilingual dialogue evaluators, this paper introduces GenResCoh (Generated Responses targeting Coherence).
-
31 Mar 2024 1 repository listedDue to the scarcity of annotated data, data augmentation is commonly used for training coherence evaluation models.
-
15 Feb 2024 1 repository listedRecent large language models (LLMs) have shown remarkable performance in aligning generated text with user intentions across various tasks.
-
18 Dec 2023 1 repository listedProviding natural language explanations for recommendations is particularly useful from the perspective of a non-expert user.
-
20 Jun 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedWe investigate CDM for open-domain text generation evaluation under two paradigms: 1) _Generative_ CDM, which harnesses the contrast of two language models' distributions to generate synthetic examples for training…
-
4 Feb 2023 1 repository listedExperiments also show that the generated captions are more coherent than that of baselines according to caption entity scores, caption Rouge scores, the two proposed coherence evaluation metrics, and human evaluations.
-
25 May 2022 1 repository listedDecision making theories such as Fuzzy-Trace Theory (FTT) suggest that individuals tend to rely on gist, or bottom-line meaning, in the text when making decisions.
-
19 May 2022 1 repository listedIn this work, we introduce SNaC, a narrative coherence evaluation framework rooted in fine-grained annotations for long summaries.
-
18 Mar 2022 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)We also show that DEAM can distinguish between coherent and incoherent dialogues generated by baseline manipulations, whereas those baseline models cannot detect incoherent examples generated by DEAM.
-
1 Nov 2021 1 repository listedEntity grids and entity graphs are two frameworks for modeling local coherence.
-
1 Jun 2021 1 repository listedTo address these limitations, we propose Quantifiable Dialogue Coherence Evaluation (QuantiDCE), a novel framework aiming to train a quantifiable dialogue coherence metric that can reflect the actual human rating…
-
7 May 2021 1 repository listedCoherent discourse is distinguished from a mere collection of utterances by the satisfaction of a diverse set of constraints, for example choice of expression, logical relation between denoted events, and implicit…
-
30 Sep 2020 1 repository listedWe evaluate our approach on the Grammarly Corpus for Discourse Coherence (GCDC) and show that when ensembled with the current state of the art, we can achieve the new state of the art accuracy on this benchmark.
-
14 May 2018 1 repository listedTo date there has been very little work on assessing discourse coherence methods on real-world data.
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections