Papers › Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling

Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling

27 Oct 2022arXiv:2210.15231archive 2025-07-28

Peijie Jiang, Dingkun Long, Yanzhao Zhang, Pengjun Xie, Meishan Zhang, Min Zhang

Boundary information is critical for various Chinese language processing tasks, such as word segmentation, part-of-speech tagging, and named entity recognition. Previous studies usually resorted to the use of a high-quality external lexicon, where lexicon items can offer explicit boundary information. However, to ensure the quality of the lexicon, great human effort is always necessary, which has been generally ignored. In this work, we suggest unsupervised statistical boundary information instead, and propose an architecture to encode the information directly into pre-trained language models, resulting in Boundary-Aware BERT (BABERT). We apply BABERT for feature induction of Chinese sequence labeling tasks. Experimental results on ten benchmarks of Chinese sequence labeling demonstrate that BABERT can provide consistent improvements on all datasets. In addition, our method can complement previous supervised lexicon exploration, where further improvements can be achieved when integrated with external lexicon information.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

modelscope/AdaSeq officialpytorch report
modelscope/modelscope pytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Chinese Named Entity RecognitionChinese Word SegmentationLanguage ModelingLanguage ModellingNamed Entity RecognitionNamed Entity Recognition (NER)Part-Of-Speech Tagging

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Chinese Word Segmentation CTB6 BABERT-LE F1 97.56 #2 of 4 Archive leaderboard report
Chinese Word Segmentation CTB6 BABERT F1 97.45 #3 of 4 Archive leaderboard report
Chinese Word Segmentation MSR BABERT-LE F1 98.63 #1 of 6 Archive leaderboard report
Chinese Word Segmentation MSR BABERT F1 98.44 #2 of 6 Archive leaderboard report
Chinese Word Segmentation MSRA BABERT-LE F1 98.63 #1 of 3 Archive leaderboard report
Chinese Word Segmentation MSRA BABERT F1 98.44 #2 of 3 Archive leaderboard report
Chinese Word Segmentation PKU BABERT-LE F1 96.84 #1 of 4 Archive leaderboard report
Chinese Word Segmentation PKU BABERT F1 96.70 #3 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections