Papers › StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding

StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding

13 Aug 2019ICLR 2020 1arXiv:1908.04577archive 2025-07-28

Wei Wang, Bin Bi, Ming Yan, Chen Wu, Zuyi Bao, Jiangnan Xia, Liwei Peng, Luo Si

Recently, the pre-trained language model, BERT (and its robustly optimized version RoBERTa), has attracted a lot of attention in natural language understanding (NLU), and achieved state-of-the-art accuracy in various NLU tasks, such as sentiment classification, natural language inference, semantic textual similarity and question answering. Inspired by the linearization exploration work of Elman [8], we extend BERT to a new model, StructBERT, by incorporating language structures into pre-training. Specifically, we pre-train StructBERT with two auxiliary tasks to make the most of the sequential order of words and sentences, which leverage language structures at the word and sentence levels, respectively. As a result, the new model is adapted to different levels of language understanding required by downstream tasks. The StructBERT with structural pre-training gives surprisingly good empirical results on a variety of downstream tasks, including pushing the state-of-the-art on the GLUE benchmark to 89.0 (outperforming all published models), the F1 score on SQuAD v1.1 question answering to 93.0, the accuracy on SNLI to 91.7.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingLinguistic AcceptabilityNatural Language InferenceNatural Language UnderstandingParaphrase IdentificationQuestion AnsweringSemantic Textual SimilaritySentenceSentiment AnalysisSentiment Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Linguistic Acceptability CoLA StructBERTRoBERTa ensemble Accuracy 69.2% #13 of 43 Archive leaderboard report
Natural Language Inference MultiNLI Adv-RoBERTa ensemble Matched 91.1 #8 of 67 Archive leaderboard report
Natural Language Inference MultiNLI Adv-RoBERTa ensemble Mismatched 90.7 #8 of 67 Archive leaderboard report
Natural Language Inference QNLI StructBERTRoBERTa ensemble Accuracy 99.2% #2 of 43 Archive leaderboard report
Natural Language Inference RTE Adv-RoBERTa ensemble Accuracy 88.7% #17 of 90 Archive leaderboard report
Natural Language Inference WNLI StructBERTRoBERTa ensemble Accuracy 89.7 #7 of 23 Archive leaderboard report
Paraphrase Identification Quora Question Pairs StructBERTRoBERTa ensemble Accuracy 90.7 #6 of 31 Archive leaderboard report
Paraphrase Identification Quora Question Pairs StructBERTRoBERTa ensemble F1 74.4 #6 of 31 Archive leaderboard report
Paraphrase Identification WikiHop StructBERTRoBERTa ensemble Accuracy 90.7% #1 of 1 Archive leaderboard report
Semantic Textual Similarity MRPC StructBERTRoBERTa ensemble Accuracy 91.5% #4 of 45 Archive leaderboard report
Semantic Textual Similarity MRPC StructBERTRoBERTa ensemble F1 93.6% #4 of 45 Archive leaderboard report
Semantic Textual Similarity STS Benchmark StructBERTRoBERTa ensemble Pearson Correlation 0.928 #2 of 66 Archive leaderboard report
Semantic Textual Similarity STS Benchmark StructBERTRoBERTa ensemble Spearman Correlation 0.924 #2 of 66 Archive leaderboard report
Sentiment Analysis SST-2 Binary classification StructBERTRoBERTa ensemble Accuracy 97.1 #6 of 87 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections