Papers › BenchIE: A Framework for Multi-Faceted Fact-Based Open Information Extraction Evaluation

BenchIE: A Framework for Multi-Faceted Fact-Based Open Information Extraction Evaluation

14 Sep 2021ACL 2022 5arXiv:2109.06850archive 2025-07-28

Kiril Gashteovski, Mingying Yu, Bhushan Kotnis, Carolin Lawrence, Mathias Niepert, Goran Glavaš

Intrinsic evaluations of OIE systems are carried out either manually -- with human evaluators judging the correctness of extractions -- or automatically, on standardized benchmarks. The latter, while much more cost-effective, is less reliable, primarily because of the incompleteness of the existing OIE benchmarks: the ground truth extractions do not include all acceptable variants of the same fact, leading to unreliable assessment of the models' performance. Moreover, the existing OIE benchmarks are available for English only. In this work, we introduce BenchIE: a benchmark and evaluation framework for comprehensive evaluation of OIE systems for English, Chinese, and German. In contrast to existing OIE benchmarks, BenchIE is fact-based, i.e., it takes into account informational equivalence of extractions: our gold standard consists of fact synsets, clusters in which we exhaustively list all acceptable surface forms of the same fact. Moreover, having in mind common downstream applications for OIE, we make BenchIE multi-faceted; i.e., we create benchmark variants that focus on different facets of OIE evaluation, e.g., compactness or minimality of extractions. We benchmark several state-of-the-art OIE systems using BenchIE and demonstrate that these systems are significantly less effective than indicated by existing OIE benchmarks. We make BenchIE (data and evaluation code) publicly available on https://github.com/gkiril/benchie.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

gkiril/benchie officialmentioned in papermentioned on GitHubNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Open Information Extraction

Datasets

Introduced by this paper, per the archive.

BenchIE

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Open Information Extraction BenchIE ClausIE F1 0.34 #1 of 11 Archive leaderboard report
Open Information Extraction BenchIE ClausIE Precision 0.50 #1 of 11 Archive leaderboard report
Open Information Extraction BenchIE ClausIE Recall 0.26 #1 of 11 Archive leaderboard report
Open Information Extraction BenchIE MinIE Precision 0.43 #2 of 11 Archive leaderboard report
Open Information Extraction BenchIE MinIE Recall 0.28 #2 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (EN) F1 0.23 #4 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (EN) Precision 0.39 #4 of 11 Archive leaderboard report
Open Information Extraction BenchIE ROIE-T F1 0.13 #5 of 11 Archive leaderboard report
Open Information Extraction BenchIE ROIE-T Precision 0.37 #5 of 11 Archive leaderboard report
Open Information Extraction BenchIE ROIE-T Recall 0.08 #5 of 11 Archive leaderboard report
Open Information Extraction BenchIE OpenIE6 F1 0.25 #6 of 11 Archive leaderboard report
Open Information Extraction BenchIE OpenIE6 Precision 0.31 #6 of 11 Archive leaderboard report
Open Information Extraction BenchIE OpenIE6 Recall 0.21 #6 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (ZH) F1 0.17 #7 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (ZH) Precision 0.26 #7 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (ZH) Recall 0.13 #7 of 11 Archive leaderboard report
Open Information Extraction BenchIE ROIE-N F1 0.13 #8 of 11 Archive leaderboard report
Open Information Extraction BenchIE ROIE-N Precision 0.20 #8 of 11 Archive leaderboard report
Open Information Extraction BenchIE ROIE-N Recall 0.09 #8 of 11 Archive leaderboard report
Open Information Extraction BenchIE Stanford OIE F1 0.13 #9 of 11 Archive leaderboard report
Open Information Extraction BenchIE Stanford OIE Precision 0.11 #9 of 11 Archive leaderboard report
Open Information Extraction BenchIE Stanford OIE Recall 0.16 #9 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (DE) F1 0.04 #10 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (DE) Precision 0.09 #10 of 11 Archive leaderboard report
Open Information Extraction BenchIE M2OIE (DE) Recall 0.03 #10 of 11 Archive leaderboard report
Open Information Extraction BenchIE Naive OIE F1 0.03 #11 of 11 Archive leaderboard report
Open Information Extraction BenchIE Naive OIE Precision 0.03 #11 of 11 Archive leaderboard report
Open Information Extraction BenchIE Naive OIE Recall 0.02 #11 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections