{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-danish-gigaword-corpus","title":"The Danish Gigaword Corpus","arxiv_id":null,"date":"2021-05-01","proceeding":"NoDaLiDa 2021 5","authors":["Leon Strømberg-Derczynski","Manuel Ciosici","Rebekah Baglini","Morten H. Christiansen","Jacob Aarup Dalsgaard","Riccardo Fusaroli","Peter Juel Henrichsen","Rasmus Hvingelby","Andreas Kirkedal","Alex Speed Kjeldsen","Claus Ladefoged","Finn Årup Nielsen","Jens Madsen","Malte Lau Petersen","Jonathan Hvithamar Rystrøm","Daniel Varab"],"abstract":"Danish language technology has been hindered by a lack of broad-coverage corpora at the scale modern NLP prefers. This paper describes the Danish Gigaword Corpus, the result of a focused effort to provide a diverse and freely-available one billion word corpus of Danish text. The Danish Gigaword corpus covers a wide array of time periods, domains, speakers’ socio-economic status, and Danish dialects.","url_abs":"https://aclanthology.org/2021.nodalida-main.46","url_pdf":"https://aclanthology.org/2021.nodalida-main.46.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"dagw","name":"DAGW","full_name":"Danish Gigaword"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}