{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dbpedia-nif-open-large-scale-and-multilingual","title":"DBpedia NIF: Open, Large-Scale and Multilingual Knowledge Extraction Corpus","arxiv_id":"1812.10315","date":"2018-12-26","proceeding":null,"authors":["Milan Dojchinovski","Julio Hernandez","Markus Ackermann","Amit Kirschenbaum","Sebastian Hellmann"],"abstract":"In the past decade, the DBpedia community has put significant amount of\neffort on developing technical infrastructure and methods for efficient\nextraction of structured information from Wikipedia. These efforts have been\nprimarily focused on harvesting, refinement and publishing semi-structured\ninformation found in Wikipedia articles, such as information from infoboxes,\ncategorization information, images, wikilinks and citations. Nevertheless,\nstill vast amount of valuable information is contained in the unstructured\nWikipedia article texts. In this paper, we present DBpedia NIF - a large-scale\nand multilingual knowledge extraction corpus. The aim of the dataset is\ntwo-fold: to dramatically broaden and deepen the amount of structured\ninformation in DBpedia, and to provide large-scale and multilingual language\nresource for development of various NLP and IR task. The dataset provides the\ncontent of all articles for 128 Wikipedia languages. We describe the dataset\ncreation process and the NLP Interchange Format (NIF) used to model the\ncontent, links and the structure the information of the Wikipedia articles. The\ndataset has been further enriched with about 25% more links and selected\npartitions published as Linked Data. Finally, we describe the maintenance and\nsustainability plans, and selected use cases of the dataset from the TextExt\nknowledge extraction challenge.","url_abs":"http://arxiv.org/abs/1812.10315v1","url_pdf":"http://arxiv.org/pdf/1812.10315v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"articles","task_name":"Articles"}],"methods":[],"datasets_introduced":[{"slug":"dbpedia-nif","name":"DBpedia NIF","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}