{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/opiec-an-open-information-extraction-corpus","title":"OPIEC: An Open Information Extraction Corpus","arxiv_id":"1904.12324","date":"2019-04-28","proceeding":"AKBC 2019","authors":["Kiril Gashteovski","Sebastian Wanner","Sven Hertling","Samuel Broscheit","Rainer Gemulla"],"abstract":"Open information extraction (OIE) systems extract relations and their\narguments from natural language text in an unsupervised manner. The resulting\nextractions are a valuable resource for downstream tasks such as knowledge base\nconstruction, open question answering, or event schema induction. In this\npaper, we release, describe, and analyze an OIE corpus called OPIEC, which was\nextracted from the text of English Wikipedia. OPIEC complements the available\nOIE resources: It is the largest OIE corpus publicly available to date (over\n340M triples) and contains valuable metadata such as provenance information,\nconfidence scores, linguistic annotations, and semantic annotations including\nspatial and temporal information. We analyze the OPIEC corpus by comparing its\ncontent with knowledge bases such as DBpedia or YAGO, which are also based on\nWikipedia. We found that most of the facts between entities present in OPIEC\ncannot be found in DBpedia and/or YAGO, that OIE facts often differ in the\nlevel of specificity compared to knowledge base facts, and that OIE open\nrelations are generally highly polysemous. We believe that the OPIEC corpus is\na valuable resource for future research on automated knowledge base\nconstruction.","url_abs":"http://arxiv.org/abs/1904.12324v1","url_pdf":"http://arxiv.org/pdf/1904.12324v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"opiec-an-open-information-extraction-corpus","repo_url":"https://github.com/uma-pi1/OPIEC","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"opiec-an-open-information-extraction-corpus","repo_url":"https://github.com/uma-pi1/OPIEC-pipeline","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"opiec-an-open-information-extraction-corpus","repo_url":"https://github.com/uma-pi1/minie","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"knowledge-base-construction","task_name":"Knowledge Base Construction"},{"task_slug":"open-information-extraction","task_name":"Open Information Extraction"},{"task_slug":"open-question","task_name":"Open-Ended Question Answering"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"specificity","task_name":"Specificity"}],"methods":[],"datasets_introduced":[{"slug":"opiec","name":"OPIEC","full_name":"Open Information Extraction Corpus"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.12324","atlas_url":"https://app.syntology.ai/?focus=1904.12324","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}