{"url":"/dataset/wikipedia-knowledge-graph-dataset","name":"Wikipedia Knowledge Graph dataset","full_name":null,"description_markdown":"Wikipedia is the largest and most read online free encyclopedia currently existing. As such, Wikipedia offers a large amount of data on all its own contents and interactions around them,  as well as different types of open data sources. This makes Wikipedia a unique data source that can be analyzed with quantitative data science techniques. However, the enormous amount of data makes it difficult to have an overview, and sometimes many of the analytical possibilities that Wikipedia offers remain unknown. In order to reduce the complexity of identifying and collecting data on Wikipedia and expanding its analytical potential, after collecting different data from various sources and processing them, we have generated a dedicated Wikipedia Knowledge Graph aimed at facilitating the analysis, contextualization of the activity and relations of Wikipedia pages, in this case limited to its English edition. We share this Knowledge Graph dataset in an open way, aiming to be useful for a wide range of researchers, such as informetricians, sociologists or data scientists.\r\n\r\nThere are a total of 9 files, all of them in tsv format, and they have been built under a relational structure. The main one that acts as the core of the dataset is the page file, after it there are 4 files with different entities related to the Wikipedia pages (category, url, pub and page_property files) and 4 other files that act as \"intermediate tables\" making it possible to connect the pages both with the latter and between pages (page_category, page_url, page_pub and page_link files).","description_withheld":null,"homepage":"https://doi.org/10.5281/zenodo.6346900","introduced_date":"2022-10-25","introduced_date_note":null,"introduced_by":{"paper":"/paper/wikinformetrics-construction-and-description","title":"Wikinformetrics: Construction and description of an open Wikipedia knowledge graph dataset for informetric purposes","first_author":null,"url":null},"license":{"name":"Creative Commons Zero v1.0 Universal","url":"https://creativecommons.org/publicdomain/zero/1.0/legalcode"},"modalities":[{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Wikipedia Knowledge Graph dataset"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}