Papers › Discourse Coherence in the Wild: A Dataset, Evaluation and Methods

Discourse Coherence in the Wild: A Dataset, Evaluation and Methods

14 May 2018WS 2018 7arXiv:1805.04993archive 2025-07-28

Alice Lai, Joel Tetreault

To date there has been very little work on assessing discourse coherence methods on real-world data. To address this, we present a new corpus of real-world texts (GCDC) as well as the first large-scale evaluation of leading discourse coherence algorithms. We show that neural models, including two that we introduce here (SentAvg and ParSeq), tend to perform best. We analyze these performance differences and discuss patterns we observed in low coherence texts in four domains.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

aylai/GCDC-corpus officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Coherence Evaluation

Datasets

Introduced by this paper, per the archive.

GCDC

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Coherence Evaluation GCDC + RST - Accuracy ParSeq Accuracy 55.09 #3 of 4 Archive leaderboard report
Coherence Evaluation GCDC + RST - F1 ParSeq Average F1 46.65 #2 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections