Papers › ERNIE: Enhanced Language Representation with Informative Entities

ERNIE: Enhanced Language Representation with Informative Entities

17 May 2019ACL 2019 7arXiv:1905.07129archive 2025-07-28

Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, Qun Liu

Neural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks. However, the existing pre-trained language models rarely consider incorporating knowledge graphs (KGs), which can provide rich structured knowledge facts for better language understanding. We argue that informative entities in KGs can enhance language representation with external knowledge. In this paper, we utilize both large-scale textual corpora and KGs to train an enhanced language representation model (ERNIE), which can take full advantage of lexical, syntactic, and knowledge information simultaneously. The experimental results have demonstrated that ERNIE achieves significant improvements on various knowledge-driven tasks, and meanwhile is comparable with the state-of-the-art model BERT on other common NLP tasks. The source code of this paper can be obtained from https://github.com/thunlp/ERNIE.

PaperPDFConference PDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="1905.07129")

Code

Syntology Ran 3 of 3 code samples harvested from 1 repository linked to this paper; 0 have no recorded run. Of those that ran: 2 ran · honoured contract; 1 ran · our draft was wrong.

By repository: official repository: 2 samples from 1 repository, 2 ran; 1 identical to code first harvested elsewhere. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

thunlp/ERNIE officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

3 samples harvested; 3 ran; 2 honoured the contract we drafted; 0 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

2ran · honoured contract
1ran · our draft was wrong

Licence: 1 of the 3 samples is pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from thunlp/ERNIE. Some samples are identical code Syntology first harvested from another repository; for those, this paper's copy is not located and its licence is not recorded. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

convert_examples_to_features thunlp/ERNIE/code/run_fewrel.py official repository ran · our draft was wrong MIT (permissive) · c248635dec6d8a67 · report
warmup_linear thunlp/ERNIE/code/run_pretrain.py official repository ran · honoured contract fingerprinted MIT (permissive) · 43676c538c76c491 · report
accuracy identical code first harvested elsewhere ran · honoured contract fingerprinted licence of this copy not recorded · eb725d5794b15f6b · report

Tasks

Entity LinkingEntity TypingKnowledge GraphsLinguistic AcceptabilityNatural Language InferenceParaphrase IdentificationRelation ClassificationRelation ExtractionSemantic Textual SimilaritySentiment Analysis

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Entity Linking FIGER ERNIE Accuracy 57.19 #1 of 1 Archive leaderboard report
Entity Linking FIGER ERNIE Macro F1 76.51 #1 of 1 Archive leaderboard report
Entity Linking FIGER ERNIE Micro F1 73.39 #1 of 1 Archive leaderboard report
Entity Typing Open Entity ERNIE F1 75.56 #3 of 3 Archive leaderboard report
Entity Typing Open Entity ERNIE Precision 78.42 #3 of 3 Archive leaderboard report
Entity Typing Open Entity ERNIE Recall 72.9 #3 of 3 Archive leaderboard report
Linguistic Acceptability CoLA ERNIE Accuracy 52.3% #35 of 43 Archive leaderboard report
Natural Language Inference MultiNLI ERNIE Matched 84.0 #34 of 67 Archive leaderboard report
Natural Language Inference MultiNLI ERNIE Mismatched 83.2 #34 of 67 Archive leaderboard report
Natural Language Inference QNLI ERNIE Accuracy 91.3% #29 of 43 Archive leaderboard report
Natural Language Inference RTE ERNIE Accuracy 68.8% #59 of 90 Archive leaderboard report
Paraphrase Identification Quora Question Pairs ERNIE F1 71.2 #15 of 31 Archive leaderboard report
Relation Classification TACRED BERT F1 66.0 #4 of 17 Archive leaderboard report
Relation Classification TACRED ERNIE F1 68.0 #6 of 17 Archive leaderboard report
Relation Extraction FewRel ERNIE F1 88.32 #2 of 2 Archive leaderboard report
Relation Extraction FewRel ERNIE Precision 88.49 #2 of 2 Archive leaderboard report
Relation Extraction FewRel ERNIE Recall 88.44 #2 of 2 Archive leaderboard report
Relation Extraction TACRED ERNIE F1 67.97 #28 of 40 Archive leaderboard report
Semantic Textual Similarity MRPC ERNIE Accuracy 88.2% #21 of 45 Archive leaderboard report
Semantic Textual Similarity STS Benchmark ERNIE Pearson Correlation 0.832 #26 of 66 Archive leaderboard report
Sentiment Analysis SST-2 Binary classification ERNIE Accuracy 93.5 #40 of 87 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutERNIELayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections