{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/paddlespeech-an-easy-to-use-all-in-one-speech-1","title":"PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit","arxiv_id":"2205.12007","date":"2022-05-20","proceeding":"NAACL (ACL) 2022 7","authors":["HUI ZHANG","Tian Yuan","Junkun Chen","Xintong Li","Renjie Zheng","Yuxin Huang","Xiaojie Chen","Enlei Gong","Zeyu Chen","Xiaoguang Hu","dianhai yu","Yanjun Ma","Liang Huang"],"abstract":"PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleSpeech to support several essential speech-to-text and text-to-speech tasks. PaddleSpeech achieves competitive or state-of-the-art performance on various speech datasets and implements the most popular methods. It also provides recipes and pretrained models to quickly reproduce the experimental results in this paper. PaddleSpeech is publicly avaiable at https://github.com/PaddlePaddle/PaddleSpeech.","url_abs":"https://arxiv.org/abs/2205.12007v1","url_pdf":"https://arxiv.org/pdf/2205.12007v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"paddlespeech-an-easy-to-use-all-in-one-speech-1","repo_url":"https://github.com/PaddlePaddle/PaddleSpeech","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"paddle","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"paddlespeech-an-easy-to-use-all-in-one-speech-1","repo_url":"https://github.com/PaddlePaddle/DeepSpeech","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"paddle","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"all","task_name":"All"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"environmental-sound-classification","task_name":"Environmental Sound Classification"},{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"},{"task_slug":"philosophy","task_name":"Philosophy"},{"task_slug":"speaker-diarization","task_name":"Speaker Diarization"},{"task_slug":"speaker-identification","task_name":"Speaker Identification"},{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"speech-to-text","task_name":"Speech-to-Text"},{"task_slug":"speech-to-text-translation","task_name":"Speech-to-Text Translation"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-synthesis","task_name":"Text-To-Speech Synthesis"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2205.12007","atlas_url":"https://app.syntology.ai/?focus=2205.12007","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}