Browse State-of-the-Art › Automatic Speech Recognition › Papers, page 10
Automatic Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 3,174 · with a code link: 677 · where Syntology ran a sample: 79 (62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (79 of 3,174 tagged: 62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 10 of 32: papers 901 to 1,000 of 3,174, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO1 Nov 2024 0 repositories listed
-
Augmenting Polish Automatic Speech Recognition System With Synthetic Data30 Oct 2024 0 repositories listed
-
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising30 Oct 2024 0 repositories listed
-
Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription29 Oct 2024 0 repositories listed
-
Asynchronous Tool Usage for Real-Time Agents28 Oct 2024 0 repositories listed
-
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs27 Oct 2024 0 repositories listed
-
A Survey on Speech Large Language Models24 Oct 2024 0 repositories listed
-
Evaluating and Improving Automatic Speech Recognition Systems for Korean Meteorological Experts24 Oct 2024 0 repositories listed
-
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams23 Oct 2024 0 repositories listed
-
DENOASR: Debiasing ASRs through Selective Denoising22 Oct 2024 0 repositories listed
-
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap22 Oct 2024 0 repositories listed
-
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models22 Oct 2024 0 repositories listed
-
Acoustic Model Optimization over Multiple Data Sources: Merging and Valuation21 Oct 2024 0 repositories listed
-
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding21 Oct 2024 0 repositories listed
-
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach19 Oct 2024 0 repositories listed
-
AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup18 Oct 2024 0 repositories listed
-
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation17 Oct 2024 0 repositories listed
-
Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR17 Oct 2024 0 repositories listed
-
Roadmap towards Superhuman Speech Understanding using Large Language Models17 Oct 2024 0 repositories listed
-
Investigation of Speaker Representation for Target-Speaker Speech Processing15 Oct 2024 0 repositories listed
-
Automatic Speech Recognition with BERT and CTC Transformers: A Review12 Oct 2024 0 repositories listed
-
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities11 Oct 2024 0 repositories listed
-
A two-stage transliteration approach to improve performance of a multilingual ASR9 Oct 2024 0 repositories listed
-
Advocating Character Error Rate for Multilingual ASR Evaluation9 Oct 2024 0 repositories listed
-
Automatic Screening for Children with Speech Disorder using Automatic Speech Recognition: Opportunities and Challenges7 Oct 2024 0 repositories listed
-
The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities5 Oct 2024 0 repositories listed
-
Team MTS @ AutoMin 2021: An Overview of Existing Summarization Approaches and Comparison to Unsupervised Summarization Techniques4 Oct 2024 0 repositories listed
-
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems3 Oct 2024 0 repositories listed
-
Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition3 Oct 2024 0 repositories listed
-
Spoken Grammar Assessment Using LLM2 Oct 2024 0 repositories listed
-
Automatic Speech Recognition for the Ika Language1 Oct 2024 0 repositories listed
-
Alignment-Free Training for Transducer-based Multi-Talker ASR30 Sep 2024 0 repositories listed
-
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems30 Sep 2024 0 repositories listed
-
Efficient Long-Form Speech Recognition for General Speech In-Context Learning29 Sep 2024 0 repositories listed
-
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility29 Sep 2024 0 repositories listed
-
A GEN AI Framework for Medical Note Generation27 Sep 2024 0 repositories listed
-
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models27 Sep 2024 0 repositories listed
-
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study26 Sep 2024 0 repositories listed
-
Deep CLAS: Deep Contextual Listen, Attend and Spell26 Sep 2024 0 repositories listed
-
Unveiling the Role of Pretraining in Direct Speech Translation26 Sep 2024 0 repositories listed
-
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not25 Sep 2024 0 repositories listed
-
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events25 Sep 2024 0 repositories listed
-
Speech Recognition Rescoring with Large Speech-Text Foundation Models25 Sep 2024 0 repositories listed
-
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM24 Sep 2024 0 repositories listed
-
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs24 Sep 2024 0 repositories listed
-
Revisiting Acoustic Features for Robust ASR24 Sep 2024 0 repositories listed
-
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices24 Sep 2024 0 repositories listed
-
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering20 Sep 2024 0 repositories listed
-
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper20 Sep 2024 0 repositories listed
-
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction20 Sep 2024 0 repositories listed
-
LM-assisted keyword biasing with Aho-Corasick algorithm for Transducer-based ASR20 Sep 2024 0 repositories listed
-
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection20 Sep 2024 0 repositories listed
-
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space19 Sep 2024 0 repositories listed
-
Personalized Speech Recognition for Children with Test-Time Adaptation19 Sep 2024 0 repositories listed
-
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR18 Sep 2024 0 repositories listed
-
Chain-of-Thought Prompting for Speech Translation17 Sep 2024 0 repositories listed
-
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text17 Sep 2024 0 repositories listed
-
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses17 Sep 2024 0 repositories listed
-
WER We Stand: Benchmarking Urdu ASR Models17 Sep 2024 0 repositories listed
-
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora17 Sep 2024 0 repositories listed
-
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models16 Sep 2024 0 repositories listed
-
An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems16 Sep 2024 0 repositories listed
-
Augmenting Automatic Speech Recognition Models with Disfluency Detection16 Sep 2024 0 repositories listed
-
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition16 Sep 2024 0 repositories listed
-
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition15 Sep 2024 0 repositories listed
-
ASR Error Correction using Large Language Models14 Sep 2024 0 repositories listed
-
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments13 Sep 2024 0 repositories listed
-
Exploring SSL Discrete Tokens for Multilingual ASR13 Sep 2024 0 repositories listed
-
Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages13 Sep 2024 0 repositories listed
-
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation13 Sep 2024 0 repositories listed
-
Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech13 Sep 2024 0 repositories listed
-
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?13 Sep 2024 0 repositories listed
-
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training13 Sep 2024 0 repositories listed
-
Full-text Error Correction for Chinese Speech Recognition with Large Language Model12 Sep 2024 0 repositories listed
-
The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language12 Sep 2024 0 repositories listed
-
Enhancing CTC-Based Visual Speech Recognition11 Sep 2024 0 repositories listed
-
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition10 Sep 2024 0 repositories listed
-
Keyword-Aware ASR Error Augmentation for Robust Dialogue State Tracking10 Sep 2024 0 repositories listed
-
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR9 Sep 2024 0 repositories listed
-
An investigation of modularity for noise robustness in conformer-based ASR9 Sep 2024 0 repositories listed
-
Evaluation of real-time transcriptions using end-to-end ASR models9 Sep 2024 0 repositories listed
-
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge9 Sep 2024 0 repositories listed
-
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge9 Sep 2024 0 repositories listed
-
Retrieval Augmented Correction of Named Entity Speech Recognition Errors9 Sep 2024 0 repositories listed
-
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection8 Sep 2024 0 repositories listed
-
Probing self-attention in self-supervised speech models for cross-linguistic differences4 Sep 2024 0 repositories listed
-
Quantification of stylistic differences in human- and ASR-produced transcripts of African American English4 Sep 2024 0 repositories listed
-
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations4 Sep 2024 0 repositories listed
-
Reassessing Noise Augmentation Methods in the Context of Adversarial Speech3 Sep 2024 0 repositories listed
-
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR3 Sep 2024 0 repositories listed
-
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka3 Sep 2024 0 repositories listed
-
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR2 Sep 2024 0 repositories listed
-
Comparing Discrete and Continuous Space LLMs for Speech Recognition1 Sep 2024 0 repositories listed
-
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition1 Sep 2024 0 repositories listed
-
Advancing Multi-talker ASR Performance with Large Language Models30 Aug 2024 0 repositories listed
-
Developing an End-to-End Framework for Predicting the Social Communication Severity Scores of Children with Autism Spectrum Disorder30 Aug 2024 0 repositories listed
-
Speaker Tagging Correction With Non-Autoregressive Language Models30 Aug 2024 0 repositories listed
-
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction29 Aug 2024 0 repositories listed
-
Literary and Colloquial Dialect Identification for Tamil using Acoustic Features27 Aug 2024 0 repositories listed
-
Automatic recognition and detection of aphasic natural speech26 Aug 2024 0 repositories listed