Browse State-of-the-Art › Sentence segmentation
Sentence segmentation
22 papers with code · 0 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
22 shown of 22 papers with code (67 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Jun 2024 3 repositories listedWe introduce a new model - Segment any Text (SaT) - to solve this problem.
-
28 Nov 2023 3 repositories listedThis study introduces Ascle, a pioneering natural language processing (NLP) toolkit designed for medical text generation.
-
30 May 2023 2 repositories listedMany NLP pipelines split text into sentences as one of the crucial preprocessing steps.
-
1 Jun 2022 2 repositories listedBecause we also collected translation process data in the form of keystroke logging, our dataset can be used as part of different research strands such as translation process research, learner corpus research, and…
-
21 Aug 2020 2 repositories listedSummarization of speech is a difficult problem due to the spontaneity of the flow, disfluencies, and other issues that are not usually encountered in written texts.
-
30 Apr 2020 2 repositories listedIn lexical semantics, full-sentence segmentation and segment labeling of various phenomena are generally treated separately, despite their interdependence.
-
23 Jan 2025 1 repository listedUsing this model, we constructed a vocabulary of DNA words and mapped DNA words to their English equivalents.
-
31 Mar 2024 1 repository listedThe texts have been enriched with seven annotation layers: (i) tokenization layer; (ii) sentence segmentation layer; (iii) lemmatization layer; (iv) morphological layer; (v) dependency layer; (vi) dependency function…
-
17 Oct 2023 1 repository listedWhile large language models (LLMs) have made considerable advancements in understanding and generating unstructured text, their application in structured data remains underexplored.
-
23 Feb 2023 1 repository listedParsing spoken dialogue presents challenges that parsing text does not, including a lack of clear sentence boundaries.
-
8 Nov 2022 1 repository listedWe present SLATE, a sequence labeling approach for extracting tasks from free-form content such as digitally handwritten (or "inked") notes on a virtual whiteboard.
-
2 Mar 2022 1 repository listedAs a solution, we present Mukayese, a set of NLP benchmarks for the Turkish language that contains several NLP tasks.
-
1 Aug 2021 1 repository listedThe sentence is a fundamental unit of text processing.
-
22 Feb 2021 1 repository listedThis paper explores the difficulties of annotating transcribed spoken Dutch-Frisian code-switch utterances into Universal Dependencies.
-
9 Jan 2021 1 repository listedFinally, we create a demo video for Trankit at: https://youtu.
-
16 Nov 2020 1 repository listedTexts obtained from web are noisy and do not necessarily follow the orthographic sentence and word boundary rules.
-
20 Sep 2020 1 repository listedWith the segmenter and the two methods combined, we compile a high-quality Bengali-English parallel corpus comprising of 2.
-
21 Aug 2020 1 repository listedSummarization of speech is a difficult problem due to the spontaneity of the flow, disfluencies, and other issues that are not usually encountered in written texts.
-
9 Apr 2020 1 repository listedThe Kurdish language is a multi-dialect, under-resourced language which is written in different scripts.
-
22 Apr 2019 1 repository listedIn this work, we argue that the task should be performed on a more fine-grained level of sequence labeling.
-
29 Jan 2019 1 repository listedThis paper describes Stanford's system at the CoNLL 2018 UD Shared Task.
-
10 Sep 2018 1 repository listedThis paper describes our submission to CoNLL 2018 UD Shared Task.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections