{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-graph-structured-dataset-for-wikipedia","title":"A Graph-structured Dataset for Wikipedia Research","arxiv_id":"1903.08597","date":"2019-03-20","proceeding":null,"authors":["Nicolas Aspert","Volodymyr Miz","Benjamin Ricaud","Pierre Vandergheynst"],"abstract":"Wikipedia is a rich and invaluable source of information. Its central place\non the Web makes it a particularly interesting object of study for scientists.\nResearchers from different domains used various complex datasets related to\nWikipedia to study language, social behavior, knowledge organization, and\nnetwork theory. While being a scientific treasure, the large size of the\ndataset hinders pre-processing and may be a challenging obstacle for potential\nnew studies. This issue is particularly acute in scientific domains where\nresearchers may not be technically and data processing savvy. On one hand, the\nsize of Wikipedia dumps is large. It makes the parsing and extraction of\nrelevant information cumbersome. On the other hand, the API is straightforward\nto use but restricted to a relatively small number of requests. The middle\nground is at the mesoscopic scale when researchers need a subset of Wikipedia\nranging from thousands to hundreds of thousands of pages but there exists no\nefficient solution at this scale.\n  In this work, we propose an efficient data structure to make requests and\naccess subnetworks of Wikipedia pages and categories. We provide convenient\ntools for accessing and filtering viewership statistics or \"pagecounts\" of\nWikipedia web pages. The dataset organization leverages principles of graph\ndatabases that allows rapid and intuitive access to subgraphs of Wikipedia\narticles and categories. The dataset and deployment guidelines are available on\nthe LTS2 website \\url{https://lts2.epfl.ch/Datasets/Wikipedia/}.","url_abs":"http://arxiv.org/abs/1903.08597v1","url_pdf":"http://arxiv.org/pdf/1903.08597v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-graph-structured-dataset-for-wikipedia","repo_url":"https://github.com/epfl-lts2/sparkwiki","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"articles","task_name":"Articles"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}