Papers › Automatic Speech Recognition in German: A Detailed Error Analysis

Automatic Speech Recognition in German: A Detailed Error Analysis

3 Aug 2022IEEE International Conference on Omni-layer Intelligent Systems (COINS) 2022 8archive 2025-07-28

Johannes Wirth, René Peinl

The amount of freely available systems for automatic speech recognition (ASR) based on neural networks is growing steadily, with equally increasingly reliable predictions. However, the evaluation of trained models is typically exclusively based on statistical metrics such as WER or CER, which do not provide any insight into the nature or impact of the errors produced when predicting transcripts from speech input. This work presents a selection of ASR model architectures that are pretrained on the German language and evaluates them on a benchmark of diverse test datasets. It identifies cross-architectural prediction errors, classifies those into categories and traces the sources of errors per category back into training data as well as other sources. Finally, it discusses solutions in order to create qualitatively better training datasets and more robust ASR systems.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Recognitionspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Automatic Speech Recognition (ASR) HUI speech corpus Conformer Transducer WER (%) 1.89% #1 of 1 Archive leaderboard report
Automatic Speech Recognition (ASR) M-AILabs speech dataset Conformer Transducer WER (%) 4.28% #1 of 1 Archive leaderboard report
Automatic Speech Recognition (ASR) The Spoken Wikipedia Corpora Conformer Transducer WER (%) 8.04% #1 of 1 Archive leaderboard report
Automatic Speech Recognition (ASR) VoxPopuli Conformer Transducer (German) WER (%) 8.98% #1 of 1 Archive leaderboard report
Automatic Speech Recognition (ASR) Voxforge German Conformer Transducer WER (%) 3.36% #1 of 1 Archive leaderboard report
Speech Recognition Common Voice German Conformer Transducer (no LM) Test WER 6.28% #6 of 14 Archive leaderboard report
Speech Recognition TUDA Conformer-Transducer (no LM) Test WER 5.82% #1 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Test

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections