Browse State-of-the-Art › Open Information Extraction
Open Information Extraction
67 papers with code · 13 benchmarks · 14 datasets archive 2025-07-28
In natural language processing, open information extraction is the task of generating a structured, machine-readable representation of the information in text, usually in the form of triples or n-ary propositions (Source: Wikipedia).
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
13 leaderboard tables shown for this task, 13 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 13 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
14 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 67 papers with code (207 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Apr 2019 3 repositories listedIn this paper, we release, describe, and analyze an OIE corpus called OPIEC, which was extracted from the text of English Wikipedia.
-
12 Aug 2023 2 repositories listedCross-lingual open information extraction aims to extract structured information from raw text across multiple languages.
-
7 Aug 2023 2 repositories listed Syntology ran 5 of 12 samples · 7 unverifiedInstruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca and Vicuna.
-
22 Jun 2022 2 repositories listedIn this paper, we propose CMVC, a novel unsupervised framework that leverages these two views of knowledge jointly for canonicalizing OKBs without the need of manually annotated labels.
-
1 Nov 2016 2 repositories listed
-
1 Jun 2016 2 repositories listed
-
28 May 2025 1 repository listedScientific research heavily depends on suitable datasets for method validation, but existing academic platforms with dataset management like PapersWithCode suffer from inefficiencies in their manual workflow.
-
18 Apr 2025 1 repository listedCompared to the baseline of unshortened (long) contexts, our experiments on four Indic languages (Hindi, Tamil, Telugu, and Urdu) demonstrate that context-shortening techniques yield an average improvement of 4\% in…
-
18 Feb 2025 1 repository listedThe effectiveness of reasoning-oriented prompting methods such as Chain-of-Thought, Reasoning and Acting, while improved with task demonstrations, does not surpass alternative methods.
-
27 Jun 2024 1 repository listedWe demonstrate the advantages of statements by applying our model to over 2700 tables from ESG reports.
-
5 Apr 2024 1 repository listedA principal issue is that, in prior methods, the KG schema has to be included in the LLM prompt to generate valid triplets; larger and more complex schemas easily exceed the LLMs' context window length.
-
16 Mar 2024 1 repository listedTo train the model, we manually annotated a large-scale Chinese OIE dataset.
-
13 Feb 2024 1 repository listedUnsupervised learning objectives like autoregressive and masked language modeling constitute a significant part in producing pre-trained representations that perform various downstream applications from natural language…
-
20 Jan 2024 1 repository listedOpen information extraction (OpenIE) aims to extract the schema-free triplets in the form of (\emph{subject}, \emph{predicate}, \emph{object}) from a given sentence.
-
23 Oct 2023 1 repository listedOpen Information Extraction (OIE) methods extract facts from natural language text in the form of ("subject"; "relation"; "object") triples.
-
22 Jun 2023 1 repository listedStructured knowledge bases (KBs) are the backbone of many know\-ledge-intensive applications, and their automated construction has received considerable attention.
-
23 May 2023 1 repository listed Syntology ran 1 of 9 samples · 8 unverifiedIn this paper, we present the first benchmark that simulates the evaluation of open information extraction models in the real world, where the syntactic and expressive distributions under the same knowledge meaning may…
-
23 May 2023 1 repository listedWe address the problem of negative transfer in TD by coupling triggers between domains using subject-object relations obtained from a rule-based open information extraction (OIE) system.
-
7 May 2023 1 repository listedWe formally define the research problem of tuple-level speculation detection and conduct a detailed data analysis on the LSOIE dataset which contains labels for speculative tuples.
-
5 May 2023 1 repository listedAccordingly, we propose a simple BERT-based model for sentence chunking, and propose Chunk-OIE for tuple extraction on top of SaC.
-
17 Jan 2023 1 repository listedIn this paper, we propose a syntactically robust training framework that enables models to be trained on a syntactic-abundant distribution based on diverse paraphrase generation.
-
5 Dec 2022 1 repository listedIn this paper, we model both constituency and dependency trees into word-level graphs, and enable neural OpenIE to learn from the syntactic structures.
-
13 Nov 2022 1 repository listedAutomated completion of open knowledge bases (Open KBs), which are constructed from triples of the form (subject phrase, relation phrase, object phrase), obtained via open information extraction (Open IE) system, are…
-
24 Jun 2022 1 repository listedOur model sets the new state of the art performance of 67.
-
21 May 2022 1 repository listed Syntology ran 7 of 13 samples · 6 unverifiedWe introduce a method for improving the structural understanding abilities of language models.
-
5 May 2022 1 repository listedOur experiments on CaRB and Wire57 datasets indicate that CompactIE finds 1.
-
1 May 2022 1 repository listedProgress with supervised Open Information Extraction (OpenIE) has been primarily limited to English due to the scarcity of training data in other languages.
-
25 Jan 2022 1 repository listedWe argue that the text and HTML structure together convey important semantics of the content and therefore warrant a special treatment for their representation learning.
-
14 Jan 2022 1 repository listedA typical information extraction pipeline consists of token- or span-level classification models coupled with a series of pre- and post-processing scripts.
-
5 Dec 2021 1 repository listedInstant analysis of cybersecurity reports is a fundamental challenge for security experts as an immeasurable amount of cyber information is generated on a daily basis, which necessitates automated information extraction…
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections