{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/kimi-audio-technical-report","title":"Kimi-Audio Technical Report","arxiv_id":"2504.18425","date":"2025-04-25","proceeding":null,"authors":["KimiTeam","Ding Ding","Zeqian Ju","Yichong Leng","Songxiang Liu","Tong Liu","Zeyu Shang","Kai Shen","Wei Song","Xu Tan","Heyi Tang","Zhengtao Wang","Chu Wei","Yifei Xin","Xinran Xu","Jianwei Yu","Yutao Zhang","Xinyu Zhou","Y. Charles","Jun Chen","Yanru Chen","Yulun Du","Weiran He","Zhenxing Hu","Guokun Lai","Qingcheng Li","Yangyang Liu","Weidong Sun","Jianzhou Wang","Yuzhi Wang","Yuefeng Wu","Yuxin Wu","Dongchao Yang","Hao Yang","Ying Yang","Zhilin Yang","Aoxiong Yin","Ruibin Yuan","Yutong Zhang","Zaida Zhou"],"abstract":"We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio.","url_abs":"https://arxiv.org/abs/2504.18425v1","url_pdf":"https://arxiv.org/pdf/2504.18425v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"kimi-audio-technical-report","repo_url":"https://github.com/moonshotai/kimi-audio","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"audio-question-answering","task_name":"Audio Question Answering"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2504.18425","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2504.18425"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/moonshotai/kimi-audio","reach":null}],"summary":{"ran_honours":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"318777afeb70ea36","entry":"compute_loss","repo":"moonshotai/kimi-audio","repo_kind":"official","path":"finetune.py","file_url":"https://github.com/moonshotai/kimi-audio/blob/HEAD/finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"318777afeb70ea36"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}