Papers › LATTE: Lattice ATTentive Encoding for Character-based Word Segmentation

LATTE: Lattice ATTentive Encoding for Character-based Word Segmentation

1 Jun 2023Journal of Natural Language Processing 2023 6archive 2025-07-28

Thodsaporn Chay-intr, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura

A character sequence comprises at least one or more segmentation alternatives. This can be considered segmentation ambiguity and may weaken segmentation performance in word segmentation. Proper handling of such ambiguity lessens ambiguous decisions on word boundaries. Previous works have achieved remarkable segmentation performance and alleviated the ambiguity problem by incorporating the lattice, owing to its ability to capture segmentation alternatives, along with graph-based and pre-trained models. However, multiple granularity information, including character and word, in a lattice that encodes with such models may not be attentively exploited. To strengthen multi-granularity representations in a lattice, we propose the Lattice ATTentive Encoding (LATTE) method for character-based word segmentation. Our model employs the lattice structure to handle segmentation alternatives and utilizes graph neural networks along with an attention mechanism to attentively extract multi-granularity representation from the lattice for complementing character representations. Our experimental results demonstrated improvements in segmentation performance on the BCCWJ, CTB6, and BEST2010 datasets in three languages, particularly Japanese, Chinese, and Thai.

PaperPDFCode

Code

tchayintr/latte-ptm-ws officialmentioned in paperpytorch report
tchayintr/latte-ws officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Chinese Word SegmentationJapanese Word SegmentationSegmentationThai Word Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Chinese Word Segmentation CTB6 LATTE (Linguistic units, lattices, PTMs, GNNs) F1 98.07 #1 of 4 Archive leaderboard report
Japanese Word Segmentation BCCWJ LATTE (Linguistic units, lattices, PTMs, GNNs) F1-score (Word) 0.9936 #1 of 3 Archive leaderboard report
Thai Word Segmentation BEST-2010 LATTE (Linguistic units, lattices, PTMs, GNNs) F1-Score 0.9907 #1 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections