{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sylber-syllabic-embedding-representation-of","title":"Sylber: Syllabic Embedding Representation of Speech from Raw Audio","arxiv_id":"2410.07168","date":"2024-10-09","proceeding":null,"authors":["Cheol Jun Cho","Nicholas Lee","Akshat Gupta","Dhruv Agarwal","Ethan Chen","Alan W Black","Gopala K. Anumanchipalli"],"abstract":"Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such structure, resulting in dense token sequences that are costly to process. To bridge this gap, we propose a new model, Sylber, that produces speech representations with clean and robust syllabic structure. Specifically, we propose a self-supervised learning (SSL) framework that bootstraps syllabic embeddings by distilling from its own initial unsupervised syllabic segmentation. This results in a highly structured representation of speech features, offering three key benefits: 1) a fast, linear-time syllable segmentation algorithm, 2) efficient syllabic tokenization with an average of 4.27 tokens per second, and 3) novel phonological units suited for efficient spoken language modeling. Our proposed segmentation method is highly robust and generalizes to out-of-domain data and unseen languages without any tuning. By training token-to-speech generative models, fully intelligible speech can be reconstructed from Sylber tokens with a significantly lower bitrate than baseline SSL tokens. This suggests that our model effectively compresses speech into a compact sequence of tokens with minimal information loss. Lastly, we demonstrate that categorical perception-a linguistic phenomenon in speech perception-emerges naturally in Sylber, making the embedding space more categorical and sparse than previous speech features and thus supporting the high efficiency of our tokenization. Together, we present a novel SSL approach for representing speech as syllables, with significant potential for efficient speech tokenization and spoken language modeling.","url_abs":"https://arxiv.org/abs/2410.07168v2","url_pdf":"https://arxiv.org/pdf/2410.07168v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sylber-syllabic-embedding-representation-of","repo_url":"https://github.com/Berkeley-Speech-Group/sylber","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"speech-tokenization","task_name":"Speech Tokenization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.07168","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.07168"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Berkeley-Speech-Group/sylber","reach":null}],"summary":{"ran":2,"ran_honours":1,"ran_fixture":1,"unverified":2},"by_repo_kind":{"official":{"samples":6,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"230dc2d1d41ecd2a","entry":"EMAModule","repo":"Berkeley-Speech-Group/sylber","repo_kind":"official","path":"sylber/model/sylber.py","file_url":"https://github.com/Berkeley-Speech-Group/sylber/blob/HEAD/sylber/model/sylber.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"230dc2d1d41ecd2a"}},{"code_sha256_prefix":"82268ab56d59fa5a","entry":"Thresholder","repo":"Berkeley-Speech-Group/sylber","repo_kind":"official","path":"sylber/model/sylber.py","file_url":"https://github.com/Berkeley-Speech-Group/sylber/blob/HEAD/sylber/model/sylber.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"82268ab56d59fa5a"}},{"code_sha256_prefix":"80c7c332ac50a7e6","entry":"cossim","repo":"Berkeley-Speech-Group/sylber","repo_kind":"official","path":"sylber/model/sylber.py","file_url":"https://github.com/Berkeley-Speech-Group/sylber/blob/HEAD/sylber/model/sylber.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"80c7c332ac50a7e6"}},{"code_sha256_prefix":"de0aac26b1069f5e","entry":"get_segment","repo":"Berkeley-Speech-Group/sylber","repo_kind":"official","path":"sylber/model/sylber.py","file_url":"https://github.com/Berkeley-Speech-Group/sylber/blob/HEAD/sylber/model/sylber.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"de0aac26b1069f5e"}},{"code_sha256_prefix":"283277b1cc3b26b5","entry":"NoiseMixer","repo":"Berkeley-Speech-Group/sylber","repo_kind":"official","path":"sylber/model/sylber.py","file_url":"https://github.com/Berkeley-Speech-Group/sylber/blob/HEAD/sylber/model/sylber.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"283277b1cc3b26b5"}},{"code_sha256_prefix":"7792f4007897767d","entry":"Sylber","repo":"Berkeley-Speech-Group/sylber","repo_kind":"official","path":"sylber/model/sylber.py","file_url":"https://github.com/Berkeley-Speech-Group/sylber/blob/HEAD/sylber/model/sylber.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7792f4007897767d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}