Papers › Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling
Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling
Peijie Jiang, Dingkun Long, Yanzhao Zhang, Pengjun Xie, Meishan Zhang, Min Zhang
Boundary information is critical for various Chinese language processing tasks, such as word segmentation, part-of-speech tagging, and named entity recognition. Previous studies usually resorted to the use of a high-quality external lexicon, where lexicon items can offer explicit boundary information. However, to ensure the quality of the lexicon, great human effort is always necessary, which has been generally ignored. In this work, we suggest unsupervised statistical boundary information instead, and propose an architecture to encode the information directly into pre-trained language models, resulting in Boundary-Aware BERT (BABERT). We apply BABERT for feature induction of Chinese sequence labeling tasks. Experimental results on ten benchmarks of Chinese sequence labeling demonstrate that BABERT can provide consistent improvements on all datasets. In addition, our method can complement previous supervised lexicon exploration, where further improvements can be achieved when integrated with external lexicon information.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Chinese Word Segmentation | CTB6 | BABERT-LE | F1 | 97.56 | #2 of 4 | Archive leaderboard | report |
| Chinese Word Segmentation | CTB6 | BABERT | F1 | 97.45 | #3 of 4 | Archive leaderboard | report |
| Chinese Word Segmentation | MSR | BABERT-LE | F1 | 98.63 | #1 of 6 | Archive leaderboard | report |
| Chinese Word Segmentation | MSR | BABERT | F1 | 98.44 | #2 of 6 | Archive leaderboard | report |
| Chinese Word Segmentation | MSRA | BABERT-LE | F1 | 98.63 | #1 of 3 | Archive leaderboard | report |
| Chinese Word Segmentation | MSRA | BABERT | F1 | 98.44 | #2 of 3 | Archive leaderboard | report |
| Chinese Word Segmentation | PKU | BABERT-LE | F1 | 96.84 | #1 of 4 | Archive leaderboard | report |
| Chinese Word Segmentation | PKU | BABERT | F1 | 96.70 | #3 of 4 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections