{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-lingual-and-cross-domain-discourse-1","title":"Cross-lingual and cross-domain discourse segmentation of entire documents","arxiv_id":"1704.04100","date":"2017-04-13","proceeding":null,"authors":["Chloé Braud","Ophélie Lacroix","Anders Søgaard"],"abstract":"Discourse segmentation is a crucial step in building end-to-end discourse\nparsers. However, discourse segmenters only exist for a few languages and\ndomains. Typically they only detect intra-sentential segment boundaries,\nassuming gold standard sentence and token segmentation, and relying on\nhigh-quality syntactic parses and rich heuristics that are not generally\navailable across languages and domains. In this paper, we propose statistical\ndiscourse segmenters for five languages and three domains that do not rely on\ngold pre-annotations. We also consider the problem of learning discourse\nsegmenters when no labeled data is available for a language. Our fully\nsupervised system obtains 89.5% F1 for English newswire, with slight drops in\nperformance on other domains, and we report supervised and unsupervised\n(cross-lingual) results for five languages in total.","url_abs":"http://arxiv.org/abs/1704.04100v2","url_pdf":"http://arxiv.org/pdf/1704.04100v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cross-lingual-and-cross-domain-discourse-1","repo_url":"https://bitbucket.org/chloebt/discourse","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"discourse-segmentation","task_name":"Discourse Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}