Browse State-of-the-Art › Phoneme Recognition
Phoneme Recognition
27 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
27 shown of 27 papers with code (104 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Sep 2016 62 repositories listed Syntology ran 41 of 103 samples · 62 unverified · 25 pointer-only (licence)This paper introduces WaveNet, a deep neural network for generating raw audio waveforms.
-
24 Jun 2015 14 repositories listedRecurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image…
-
14 Nov 2012 7 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedOne of the key challenges in sequence transduction is learning to represent both the input and output sequences in a way that is invariant to sequential distortions such as shrinking, stretching and translating.
-
22 Mar 2013 5 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedRecurrent neural networks (RNNs) are a powerful model for sequential data.
-
29 Mar 2022 2 repositories listedWe show that fine-tuning with pseudo labels achieves a 5.
-
23 Sep 2021 2 repositories listedRecent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data.
-
21 Dec 2013 2 repositories listedCurrently, deep neural networks are the state of the art on problems such as speech recognition and computer vision.
-
13 Jun 2024 1 repository listedExperiments are conducted with HuBERT and WavLM models and evaluated on the SUPERB benchmark for two content-related tasks: automatic speech recognition (ASR) and phoneme recognition (PR).
-
10 Mar 2024 1 repository listedThese task-specific representations are used for robust performance on various downstream tasks by fine-tuning on the labelled data.
-
19 Sep 2023 1 repository listedIn the first step, the pre-trained SSL model is fine-tuned on a phoneme recognition task to obtain better representations for the pronounced phonemes.
-
7 Jun 2023 1 repository listedThis paper proposes Allophant, a multilingual phoneme recognizer.
-
5 Sep 2022 1 repository listedThis open source software tool, released under MIT license, is developed as a one-stop solution to handle different speech related text processing tasks for automatic speech recognition, text to speech synthesis and…
-
1 Jul 2022 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedOur method reduces the model to 23.
-
24 Jun 2022 1 repository listedCompared with MFCC, in the within-language scenario, the performance of these SSL speech pre-trained models on AF probing tasks achieved a maximum relative increase of 34.
-
15 Jun 2022 1 repository listedIn the field of assessing the pronunciation quality of constrained speech, the given transcriptions can play the role of a teacher.
-
1 Apr 2022 1 repository listedMost of the research on data-driven speech representation learning has focused on raw audios in an end-to-end manner, paying little attention to their internal phonological or gestural structure.
-
22 Feb 2022 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedStochastic latent variable models (LVMs) achieve state-of-the-art performance on natural image generation but are still inferior to deterministic models on speech.
-
21 Feb 2022 1 repository listedWhen trained on 36 Spanish phonemes and tested on real recordings of collaborative learning environments, a 0.
-
16 Nov 2021 1 repository listedPhonemes are defined by their relationship to words: changing a phoneme changes the word.
-
31 May 2021 1 repository listedExtensive works have tackled Language Identification (LID) in the speech domain, however their application to the singing voice trails and performances on Singing Language Identification (SLID) can be improved…
-
25 Mar 2021 1 repository listedWhile speech recognition has seen a surge in interest and research over the last decade, most machine learning models for speech recognition either require large training datasets or lots of storage and memory.
-
8 Aug 2020 1 repository listedMeasuring the performance of automatic speech recognition (ASR) systems requires manually transcribed data in order to compute the word error rate (WER), which is often time-consuming and expensive.
-
17 Dec 2019 1 repository listedAt the same time, in order to solve the problem of overfitting in the 61 phoneme recognition model on TIMIT dataset, we propose a new training method.
-
20 Jun 2018 1 repository listedQuaternion numbers and quaternion neural networks have shown their efficiency to process multidimensional inputs as entities, to encode internal dependencies, and to solve many tasks with less learning parameters than…
-
10 Jan 2017 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedMeanwhile, Connectionist Temporal Classification (CTC) with Recurrent Neural Networks (RNNs), which is proposed for labeling unsegmented sequences, makes it feasible to train an end-to-end speech recognition system…
-
26 Nov 2015 1 repository listedWe stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms.
-
7 Dec 2013 1 repository listedMost phoneme recognition state-of-the-art systems rely on a classical neural network classifiers, fed with highly tuned features, such as MFCC or PLP features.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections