{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-gutenberg-dialogue-dataset","title":"The Gutenberg Dialogue Dataset","arxiv_id":"2004.12752","date":"2020-04-27","proceeding":"EACL 2021 2","authors":["Richard Csaky","Gabor Recski"],"abstract":"Large datasets are essential for neural modeling of many NLP tasks. Current publicly available open-domain dialogue datasets offer a trade-off between quality (e.g., DailyDialog) and size (e.g., Opensubtitles). We narrow this gap by building a high-quality dataset of 14.8M utterances in English, and smaller datasets in German, Dutch, Spanish, Portuguese, Italian, and Hungarian. We extract and process dialogues from public-domain books made available by Project Gutenberg. We describe our dialogue extraction pipeline, analyze the effects of the various heuristics used, and present an error analysis of extracted dialogues. Finally, we conduct experiments showing that better response quality can be achieved in zero-shot and finetuning settings by training on our data than on the larger but much noisier Opensubtitles dataset. Our open-source pipeline (https://github.com/ricsinaruto/gutenberg-dialog) can be extended to further languages with little additional effort. Researchers can also build their versions of existing datasets by adjusting various trade-off parameters. We also built a web demo for interacting with our models: https://ricsinaruto.github.io/chatbot.html.","url_abs":"https://arxiv.org/abs/2004.12752v2","url_pdf":"https://arxiv.org/pdf/2004.12752v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-gutenberg-dialogue-dataset","repo_url":"https://github.com/ricsinaruto/gutenberg-dialog","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"gutenberg-dialog-dataset","name":"Gutenberg Dialog Dataset","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2004.12752","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2004.12752"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ricsinaruto/gutenberg-dialog","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_fixture":1,"unverified":2},"by_repo_kind":{"official":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a07f6a302fe3b9eb","entry":"top_filtering","repo":"ricsinaruto/gutenberg-dialog","repo_kind":"official","path":"gpt2_trainings_scripts/interact.py","file_url":"https://github.com/ricsinaruto/gutenberg-dialog/blob/HEAD/gpt2_trainings_scripts/interact.py","link_basis":"plan_row","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a07f6a302fe3b9eb"}},{"code_sha256_prefix":"10f1aa81cc87b047","entry":"make_logdir","repo":"ricsinaruto/gutenberg-dialog","repo_kind":"official","path":"gpt2_trainings_scripts/utils.py","file_url":"https://github.com/ricsinaruto/gutenberg-dialog/blob/HEAD/gpt2_trainings_scripts/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"10f1aa81cc87b047"}},{"code_sha256_prefix":"b7c5b442374fd09e","entry":"processed","repo":"ricsinaruto/gutenberg-dialog","repo_kind":"official","path":"gpt2_trainings_scripts/interact.py","file_url":"https://github.com/ricsinaruto/gutenberg-dialog/blob/HEAD/gpt2_trainings_scripts/interact.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b7c5b442374fd09e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}