{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-dataset-for-web-scale-knowledge-base","title":"A Dataset for Web-Scale Knowledge Base Population","arxiv_id":null,"date":"2018-06-03","proceeding":"European Semantic Web Conference 2018 6","authors":["Michael Glass","Alfio Gliozzo"],"abstract":"For many domains, structured knowledge is in short supply,\r\nwhile unstructured text is plentiful. Knowledge Base Population (KBP)\r\nis the task of building or extending a knowledge base from text, and\r\nsystems for KBP have grown in capability and scope. However, existing\r\ndatasets for KBP are all limited by multiple issues: small in size, not\r\nopen or accessible, only capable of benchmarking a fraction of the KBP\r\nprocess, or only suitable for extracting knowledge from title-oriented doc\u0002uments (documents that describe a particular entity, such as Wikipedia\r\npages). We introduce and release CC-DBP, a web-scale dataset for train\u0002ing and benchmarking KBP systems. The dataset is based on Common\r\nCrawl as the corpus and DBpedia as the target knowledge base. Criti\u0002cally, by releasing the tools to build the dataset, we enable the dataset\r\nto remain current as new crawls and DBpedia dumps are released. Also,\r\nthe modularity of the released tool set resolves a crucial tension between\r\nthe ease that a dataset can be used for a particular subtask in KBP and\r\nthe number of different subtasks it can be used to train or benchmark.","url_abs":"https://link.springer.com/chapter/10.1007/978-3-319-93417-4_17","url_pdf":"https://link.springer.com/content/pdf/10.1007/978-3-319-93417-4_17.pdf?pdf=inline%20link","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-dataset-for-web-scale-knowledge-base","repo_url":"https://github.com/IBM/cc-dbp","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"knowledge-base-population","task_name":"Knowledge Base Population"}],"methods":[],"datasets_introduced":[{"slug":"cc-dbp","name":"CC-DBP","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}