Browse State-of-the-Art › Automatic Lyrics Transcription
Automatic Lyrics Transcription
9 papers with code · 5 benchmarks · 1 dataset archive 2025-07-28
Automatic Lyrics Transcription is the task of transcribing singing voice from audio into text.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Jam-ALT English (18 rows) | AudioShake v3 | Lyrics Transcription for Humans: A Readability-Aware Benchmark | code | — | Compare |
| Jam-ALT (16 rows) | AudioShake v3 | Lyrics Transcription for Humans: A Readability-Aware Benchmark | code | — | Compare |
| Jam-ALT Spanish (16 rows) | AudioShake v3 | Lyrics Transcription for Humans: A Readability-Aware Benchmark | code | — | Compare |
| Jam-ALT German (16 rows) | AudioShake v3 | Lyrics Transcription for Humans: A Readability-Aware Benchmark | code | — | Compare |
| Jam-ALT French (16 rows) | AudioShake v3 | Lyrics Transcription for Humans: A Readability-Aware Benchmark | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
9 shown of 9 papers with code (12 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Aug 2021 2 repositories listedThis paper makes several contributions to automatic lyrics transcription (ALT) research.
-
13 Jul 2020 2 repositories listedSpeech recognition is a well developed research field so that the current state of the art systems are being used in many applications in the software industry, yet as by today, there still does not exist such robust…
-
30 Jul 2024 1 repository listedWriting down lyrics for human consumption involves not only accurately capturing word sequences, but also incorporating punctuation and formatting for clarity and to convey contextual information.
-
25 Jun 2024 1 repository listedFurthermore, we demonstrate that incorporating language information significantly enhances performance.
-
23 Nov 2023 1 repository listedCurrent automatic lyrics transcription (ALT) benchmarks focus exclusively on word content and ignore the finer nuances of written lyrics including formatting and punctuation, which leads to a potential misalignment with…
-
21 Nov 2023 1 repository listedWith the use of data augmentation and source separation model, results show that the proposed method achieves a character error rate of less than 18% on a Mandarin polyphonic dataset for lyrics transcription, and a mean…
-
29 Jun 2023 1 repository listedWe introduce LyricWhiz, a robust, multilingual, and zero-shot automatic lyrics transcription method achieving state-of-the-art performance on various lyrics transcription datasets, even in challenging genres such as…
-
7 Apr 2022 1 repository listedTo improve the robustness of lyrics transcription to the background music, we propose a strategy of combining the features that emphasize the singing vocals, i.
-
21 Jun 2021 1 repository listedRecent automatic lyrics transcription (ALT) approaches focus on building stronger acoustic models or in-domain language models, while the pronunciation aspect is seldom touched upon.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections