Browse State-of-the-Art › Text Segmentation
Text Segmentation
40 papers with code · 0 benchmarks · 7 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 40 papers with code (124 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Oct 2016 3 repositories listedWe propose a novel domain-independent framework, called CoType, that runs a data-driven text segmentation algorithm to extract entity mentions, and jointly embeds entity mentions, relation mentions, text features and…
-
6 Jan 2025 2 repositories listedReinforcement learning from human feedback (RLHF) has been widely adopted to align language models (LMs) with human preference.
-
7 Oct 2021 2 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedIn this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours…
-
25 Mar 2018 2 repositories listedText segmentation, the task of dividing a document into contiguous segments based on its semantic structure, is a longstanding challenge in language understanding.
-
24 Feb 2017 2 repositories listedThe probability of a segmented sequence is calculated as the product of the probabilities of all its segments, where each segment is modeled using existing tools such as recurrent neural networks.
-
16 Feb 2025 1 repository listedThis work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations.
-
16 Oct 2024 1 repository listed Syntology ran 5 of 5 samples · 0 unverifiedWhile Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for boosting large language models (LLMs) in knowledge-intensive tasks, it often overlooks the crucial aspect of text chunking within its…
-
31 Jul 2024 1 repository listedOne challenge of the task is that the local stroke shapes of artistic text are changeable with diversity and complexity.
-
24 Jul 2024 1 repository listedIn this paper, we propose Edge-Aware Transformers, termed EAFormer, to segment texts more accurately, especially at the edge of texts.
-
6 Mar 2024 1 repository listed Syntology ran 2 of 6 samples · 4 unverified · 6 pointer-only (licence)Our empirical findings highlight (1) detecting AI-generated sentences in hybrid texts is overall a challenging task because (1.
-
31 Jan 2024 1 repository listed Syntology ran 5 of 5 samples · 0 unverifiedWe use this TS model to iteratively generate the pixel-level text labels in a semi-automatical manner, unifying labels across the four text hierarchies in the HierText dataset.
-
29 Nov 2023 1 repository listedSemi-Markov CRF has been proposed as an alternative to the traditional Linear Chain CRF for text segmentation tasks such as Named Entity Recognition (NER).
-
13 Jun 2023 1 repository listedIn each iteration, the original image, previous text removal result, and text mask are input to the network to extract the rest part of the text segments and cleaner text removal result.
-
8 Jun 2023 1 repository listedMost existing studies focused on designing attacks to evaluate the robustness of NLP models in the English language alone.
-
27 May 2023 1 repository listedThis work compares the performance of the proposed method with other state-of-the-art (SOTA) methods on DIBCO and H-DIBCO ((Handwritten) Document Image Binarization Competition) datasets.
-
29 Nov 2022 1 repository listedIn Stage-2, we use four independent generators to separately train GAN models based on the four channels on the processed patch images to extract color foreground information.
-
28 Nov 2022 1 repository listedAs phrase extraction can be regarded as a $1$D text segmentation problem, we formulate PEG as a dual detection problem and propose a novel DQ-DETR model, which introduces dual queries to probe different features from…
-
1 Nov 2022 1 repository listedTherefore, exploring the robust text feature representations on unlabeled real images by self-supervised learning is a good solution.
-
28 Oct 2022 1 repository listedThe problem is only exacerbated by a lack of segmentation in transcripts of audio/video recordings.
-
7 Mar 2022 1 repository listedSupervised attention can alleviate the above issue, but it is character category-specific, which requires extra laborious character-level bounding box annotations and would be memory-intensive when handling languages…
-
1 Nov 2021 1 repository listedDiscourse segmentation, the first step of discourse analysis, has been shown to improve results for text summarization, translation and other NLP tasks.
-
1 Aug 2021 1 repository listedThis paper describes our contribution to SemEval-2021 Task 5: Toxic Spans Detection.
-
15 Apr 2021 1 repository listed Syntology ran 6 of 8 samples · 2 unverified · 8 pointer-only (licence)Prior methods to text segmentation are mostly at token level.
-
7 Dec 2020 1 repository listedThe growing complexity of legal cases has lead to an increasing interest in legal information retrieval systems that can effectively satisfy user-specific information needs.
-
1 Dec 2020 1 repository listedIn this paper, we address the segmentation of books of hours, Latin devotional manuscripts of the late Middle Ages, that exhibit challenging issues: a complex hierarchical entangled structure, variable content, noisy…
-
27 Nov 2020 1 repository listedWe also introduce Text Refinement Network (TexRNet), a novel text segmentation approach that adapts to the unique properties of text, e.
-
14 Nov 2020 1 repository listedNatural language segmentation (NLS), or text segmentation, refers to the process of dividing written text into meaningful units.
-
9 Nov 2020 1 repository listedBooks are typically segmented into chapters and sections, representing coherent subnarratives and topics.
-
22 May 2020 1 repository listedWe formulate the problem as a sequence labelling task, and study the performance of state of the art approaches.
-
30 Apr 2020 1 repository listedDocument and discourse segmentation are two fundamental NLP tasks pertaining to breaking up text into constituents, which are commonly used to help downstream tasks such as information retrieval or text summarization.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections