Browse State-of-the-Art › Speech Tokenization
Speech Tokenization
9 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Speech tokenization is the task of representing speech signals as a sequence of discrete units. Such representations can be later used for various downstream tasks including automatic speech recognition, text-to-speech, etc. Such representation serves as the basis of Speech Language Models.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
9 shown of 9 papers with code (21 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Apr 2025 2 repositories listed Syntology ran 14 of 20 samples · 6 unverified · 20 pointer-only (licence)To our knowledge, TASTE is the first end-to-end approach that utilizes a reconstruction objective to automatically learn a text-aligned speech tokenization and embedding suitable for spoken language modeling.
-
24 May 2025 1 repository listedRecent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio…
-
21 Nov 2024 1 repository listedSpoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality.
-
19 Oct 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedRecent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis.
-
9 Oct 2024 1 repository listed Syntology ran 4 of 6 samples · 2 unverifiedTo bridge this gap, we propose a new model, Sylber, that produces speech representations with clean and robust syllabic structure.
-
5 Oct 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)For speech in particular, the high resolution of waveforms (16, 000 samples/second or more) presents a significant challenge as speech-based language models have had to use several times more tokens per word than…
-
16 Sep 2024 1 repository listedHowever, we observe that the information aggregated in the CLS token correlates more with speaker identity than with linguistic content.
-
22 Jul 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Inspired by this success, researchers have investigated various compression-based speech tokenization methods to discretize continuous speech signals, enabling the application of language modeling techniques to discrete…
-
31 Aug 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)To improve the performance of these discrete speech tokens, we present RepCodec, a novel speech representation codec for semantic speech tokenization.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections