Papers › LAST: Language Model Aware Speech Tokenization

LAST: Language Model Aware Speech Tokenization

5 Sep 2024arXiv:2409.03701archive 2025-07-28

Arnon Turetzky, Yossi Adi

Speech tokenization serves as the foundation of speech language model (LM), enabling them to perform various tasks such as spoken language modeling, text-to-speech, speech-to-text, etc. Most speech tokenizers are trained independently of the LM training process, relying on separate acoustic models and quantization methods. Following such an approach may create a mismatch between the tokenization process and its usage afterward. In this study, we propose a novel approach to training a speech tokenizer by leveraging objectives from pre-trained textual LMs. We advocate for the integration of this objective into the process of learning discrete speech representations. Our aim is to transform features from a pre-trained speech model into a new feature space that enables better clustering for speech LMs. We empirically investigate the impact of various model design choices, including speech vocabulary size and text LM size. Our results demonstrate the proposed tokenization method outperforms the evaluated baselines considering both spoken language modeling and speech-to-text. More importantly, unlike prior work, the proposed method allows the utilization of a single pre-trained LM for processing both speech and text inputs, setting it apart from conventional tokenization approaches.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingQuantizationSpeech TokenizationSpeech-to-TextText to Speechmodeltext-to-speech

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Language Modelling SALMon LAST 1.3B Background (Domain) Consistency 56.0 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 1.3B Background (Random) Consistency 61.0 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 1.3B Background Alignment 53.0 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 1.3B Gender Consistency 68.5 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 1.3B Room Consistency 62.5 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 1.3B Sentiment Alignment 53.5 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 1.3B Sentiment Consistency 65.0 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 1.3B Speaker Consistency 64.5 #2 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Background (Domain) Consistency 55.5 #3 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Background (Random) Consistency 60.5 #3 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Background Alignment 54.5 #3 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Gender Consistency 70.5 #3 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Room Consistency 61.0 #3 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Sentiment Alignment 51.5 #3 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Sentiment Consistency 64.0 #3 of 10 Archive leaderboard report
Language Modelling SALMon LAST 350M Speaker Consistency 63.0 #3 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections