{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/kt-speech-crawler-automatic-dataset","title":"KT-Speech-Crawler: Automatic Dataset Construction for Speech Recognition from YouTube Videos","arxiv_id":"1903.00216","date":"2019-03-01","proceeding":"EMNLP 2018 11","authors":["Egor Lakomkin","Sven Magg","Cornelius Weber","Stefan Wermter"],"abstract":"In this paper, we describe KT-Speech-Crawler: an approach for automatic\ndataset construction for speech recognition by crawling YouTube videos. We\noutline several filtering and post-processing steps, which extract samples that\ncan be used for training end-to-end neural speech recognition systems. In our\nexperiments, we demonstrate that a single-core version of the crawler can\nobtain around 150 hours of transcribed speech within a day, containing an\nestimated 3.5% word error rate in the transcriptions. Automatically collected\nsamples contain reading and spontaneous speech recorded in various conditions\nincluding background noise and music, distant microphone recordings, and a\nvariety of accents and reverberation. When training a deep neural network on\nspeech recognition, we observed around 40\\% word error rate reduction on the\nWall Street Journal dataset by integrating 200 hours of the collected samples\ninto the training set. The demo (http://emnlp-demo.lakomkin.me/) and the\ncrawler code (https://github.com/EgorLakomkin/KTSpeechCrawler) are publicly\navailable.","url_abs":"http://arxiv.org/abs/1903.00216v1","url_pdf":"http://arxiv.org/pdf/1903.00216v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"kt-speech-crawler-automatic-dataset","repo_url":"https://github.com/EgorLakomkin/KTSpeechCrawler","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.00216","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1903.00216"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/EgorLakomkin/KTSpeechCrawler","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b26d5ee368b6072e","entry":"select_random_sample","repo":"EgorLakomkin/KTSpeechCrawler","repo_kind":"official","path":"webdemo/server.py","file_url":"https://github.com/EgorLakomkin/KTSpeechCrawler/blob/HEAD/webdemo/server.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b26d5ee368b6072e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}