{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-standardized-project-gutenberg-corpus-for","title":"A standardized Project Gutenberg corpus for statistical analysis of natural language and quantitative linguistics","arxiv_id":"1812.08092","date":"2018-12-19","proceeding":null,"authors":["Martin Gerlach","Francesc Font-Clos"],"abstract":"The use of Project Gutenberg (PG) as a text corpus has been extremely popular\nin statistical analysis of language for more than 25 years. However, in\ncontrast to other major linguistic datasets of similar importance, no\nconsensual full version of PG exists to date. In fact, most PG studies so far\neither consider only a small number of manually selected books, leading to\npotential biased subsets, or employ vastly different pre-processing strategies\n(often specified in insufficient details), raising concerns regarding the\nreproducibility of published results. In order to address these shortcomings,\nhere we present the Standardized Project Gutenberg Corpus (SPGC), an open\nscience approach to a curated version of the complete PG data containing more\nthan 50,000 books and more than $3 \\times 10^9$ word-tokens. Using different\nsources of annotated metadata, we not only provide a broad characterization of\nthe content of PG, but also show different examples highlighting the potential\nof SPGC for investigating language variability across time, subjects, and\nauthors. We publish our methodology in detail, the code to download and process\nthe data, as well as the obtained corpus itself on 3 different levels of\ngranularity (raw text, timeseries of word tokens, and counts of words). In this\nway, we provide a reproducible, pre-processed, full-size version of Project\nGutenberg as a new scientific resource for corpus linguistics, natural language\nprocessing, and information retrieval.","url_abs":"http://arxiv.org/abs/1812.08092v1","url_pdf":"http://arxiv.org/pdf/1812.08092v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-standardized-project-gutenberg-corpus-for","repo_url":"https://github.com/pgcorpus/gutenberg","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"a-standardized-project-gutenberg-corpus-for","repo_url":"https://github.com/pgcorpus/gutenberg-analysis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"a-standardized-project-gutenberg-corpus-for","repo_url":"https://github.com/christofs/sentlens","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[{"slug":"standardized-project-gutenberg-corpus","name":"Standardized Project Gutenberg Corpus","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.08092","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}