{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/discourse-coherence-in-the-wild-a-dataset","title":"Discourse Coherence in the Wild: A Dataset, Evaluation and Methods","arxiv_id":"1805.04993","date":"2018-05-14","proceeding":"WS 2018 7","authors":["Alice Lai","Joel Tetreault"],"abstract":"To date there has been very little work on assessing discourse coherence\nmethods on real-world data. To address this, we present a new corpus of\nreal-world texts (GCDC) as well as the first large-scale evaluation of leading\ndiscourse coherence algorithms. We show that neural models, including two that\nwe introduce here (SentAvg and ParSeq), tend to perform best. We analyze these\nperformance differences and discuss patterns we observed in low coherence texts\nin four domains.","url_abs":"http://arxiv.org/abs/1805.04993v1","url_pdf":"http://arxiv.org/pdf/1805.04993v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"discourse-coherence-in-the-wild-a-dataset","repo_url":"https://github.com/aylai/GCDC-corpus","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"coherence-evaluation","task_name":"Coherence Evaluation"}],"methods":[],"datasets_introduced":[{"slug":"gcdc","name":"GCDC","full_name":"Grammarly Corpus of Discourse Coherence"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/coherence-evaluation-on-gcdc-rst-accuracy","task":"Coherence Evaluation","dataset":"GCDC + RST - Accuracy","model":"ParSeq","rank_in_archive_order":3,"of":4,"metrics":{"Accuracy":"55.09"},"uses_additional_data":false},{"leaderboard":"/sota/coherence-evaluation-on-gcdc-rst-f1","task":"Coherence Evaluation","dataset":"GCDC + RST - F1","model":"ParSeq","rank_in_archive_order":2,"of":3,"metrics":{"Average F1":"46.65"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.04993","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}