Papers › Revisiting the Negative Data of Distantly Supervised Relation Extraction

Revisiting the Negative Data of Distantly Supervised Relation Extraction

21 May 2021ACL 2021 5arXiv:2105.10158archive 2025-07-28

Chenhao Xie, Jiaqing Liang, Jingping Liu, Chengsong Huang, Wenhao Huang, Yanghua Xiao

Distantly supervision automatically generates plenty of training samples for relation extraction. However, it also incurs two major problems: noisy labels and imbalanced training data. Previous works focus more on reducing wrongly labeled relations (false positives) while few explore the missing relations that are caused by incompleteness of knowledge base (false negatives). Furthermore, the quantity of negative labels overwhelmingly surpasses the positive ones in previous problem formulations. In this paper, we first provide a thorough analysis of the above challenges caused by negative data. Next, we formulate the problem of relation extraction into as a positive unlabeled learning task to alleviate false negative problem. Thirdly, we propose a pipeline approach, dubbed \textsc{ReRe}, that performs sentence-level relation detection then subject/object extraction to achieve sample-efficient training. Experimental results show that the proposed method consistently outperforms existing approaches and remains excellent performance even learned with a large quantity of false positive samples.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

redreamality/RERE-relation-extraction officialmentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Relation ExtractionSentence

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Relation Extraction NYT10-HRL ReRe F1 73.95 #1 of 10 Archive leaderboard report
Relation Extraction NYT10-HRL ReRe (exact) F1 73.4 #2 of 10 Archive leaderboard report
Relation Extraction NYT10-HRL TPLinker Wang et al. (2020)* F1 72.45 #4 of 10 Archive leaderboard report
Relation Extraction NYT10-HRL TPLinker Wang et al. (2020)*(exact) F1 71.93 #6 of 10 Archive leaderboard report
Relation Extraction NYT10-HRL CasRel (exact) F1 70.11 #8 of 10 Archive leaderboard report
Relation Extraction NYT10-HRL HRL Takanobu et al. (2019) F1 64.4 #10 of 10 Archive leaderboard report
Relation Extraction NYT11-HRL RERE F1 56.23 #1 of 12 Archive leaderboard report
Relation Extraction NYT11-HRL ReRe (exact) F1 55.47 #3 of 12 Archive leaderboard report
Relation Extraction NYT11-HRL HRL F1 53.8 #7 of 12 Archive leaderboard report
Relation Extraction NYT21 ReRe F1 59.62 #1 of 4 Archive leaderboard report
Relation Extraction NYT21 ReRe (exact) F1 58.88 #2 of 4 Archive leaderboard report
Relation Extraction NYT21 TPLinker(exact) F1 57.33 #3 of 4 Archive leaderboard report
Relation Extraction NYT21 CasRel (exact) F1 54.78 #4 of 4 Archive leaderboard report
Relation Extraction SKE ReRe (exact) F1 87.21 #1 of 3 Archive leaderboard report
Relation Extraction SKE CasRel (exact) F1 86.45 #2 of 3 Archive leaderboard report
Relation Extraction SKE TPLinker (exact) F1 84.32 #3 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections