Browse State-of-the-Art › Automatic Speech Recognition (ASR) › Papers, page 9
Automatic Speech Recognition (ASR)
Papers archive 2025-07-28
archive papers tagged: 3,012 · with a code link: 622 · where Syntology ran a sample: 77 (64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (77 of 3,012 tagged: 64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument)
Page 9 of 31: papers 801 to 900 of 3,012, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
DENOASR: Debiasing ASRs through Selective Denoising22 Oct 2024 0 repositories listed
-
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap22 Oct 2024 0 repositories listed
-
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models22 Oct 2024 0 repositories listed
-
Acoustic Model Optimization over Multiple Data Sources: Merging and Valuation21 Oct 2024 0 repositories listed
-
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding21 Oct 2024 0 repositories listed
-
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach19 Oct 2024 0 repositories listed
-
AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup18 Oct 2024 0 repositories listed
-
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation17 Oct 2024 0 repositories listed
-
Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR17 Oct 2024 0 repositories listed
-
Roadmap towards Superhuman Speech Understanding using Large Language Models17 Oct 2024 0 repositories listed
-
Automatic Speech Recognition with BERT and CTC Transformers: A Review12 Oct 2024 0 repositories listed
-
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities11 Oct 2024 0 repositories listed
-
A two-stage transliteration approach to improve performance of a multilingual ASR9 Oct 2024 0 repositories listed
-
Advocating Character Error Rate for Multilingual ASR Evaluation9 Oct 2024 0 repositories listed
-
Automatic Screening for Children with Speech Disorder using Automatic Speech Recognition: Opportunities and Challenges7 Oct 2024 0 repositories listed
-
The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities5 Oct 2024 0 repositories listed
-
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems3 Oct 2024 0 repositories listed
-
Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition3 Oct 2024 0 repositories listed
-
Spoken Grammar Assessment Using LLM2 Oct 2024 0 repositories listed
-
Automatic Speech Recognition for the Ika Language1 Oct 2024 0 repositories listed
-
Alignment-Free Training for Transducer-based Multi-Talker ASR30 Sep 2024 0 repositories listed
-
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems30 Sep 2024 0 repositories listed
-
Efficient Long-Form Speech Recognition for General Speech In-Context Learning29 Sep 2024 0 repositories listed
-
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility29 Sep 2024 0 repositories listed
-
A GEN AI Framework for Medical Note Generation27 Sep 2024 0 repositories listed
-
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study26 Sep 2024 0 repositories listed
-
Deep CLAS: Deep Contextual Listen, Attend and Spell26 Sep 2024 0 repositories listed
-
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events25 Sep 2024 0 repositories listed
-
Speech Recognition Rescoring with Large Speech-Text Foundation Models25 Sep 2024 0 repositories listed
-
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM24 Sep 2024 0 repositories listed
-
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs24 Sep 2024 0 repositories listed
-
Revisiting Acoustic Features for Robust ASR24 Sep 2024 0 repositories listed
-
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices24 Sep 2024 0 repositories listed
-
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering20 Sep 2024 0 repositories listed
-
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper20 Sep 2024 0 repositories listed
-
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection20 Sep 2024 0 repositories listed
-
Personalized Speech Recognition for Children with Test-Time Adaptation19 Sep 2024 0 repositories listed
-
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR18 Sep 2024 0 repositories listed
-
Chain-of-Thought Prompting for Speech Translation17 Sep 2024 0 repositories listed
-
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text17 Sep 2024 0 repositories listed
-
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses17 Sep 2024 0 repositories listed
-
WER We Stand: Benchmarking Urdu ASR Models17 Sep 2024 0 repositories listed
-
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora17 Sep 2024 0 repositories listed
-
An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems16 Sep 2024 0 repositories listed
-
Augmenting Automatic Speech Recognition Models with Disfluency Detection16 Sep 2024 0 repositories listed
-
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition16 Sep 2024 0 repositories listed
-
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition15 Sep 2024 0 repositories listed
-
ASR Error Correction using Large Language Models14 Sep 2024 0 repositories listed
-
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments13 Sep 2024 0 repositories listed
-
Exploring SSL Discrete Tokens for Multilingual ASR13 Sep 2024 0 repositories listed
-
Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages13 Sep 2024 0 repositories listed
-
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation13 Sep 2024 0 repositories listed
-
Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech13 Sep 2024 0 repositories listed
-
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training13 Sep 2024 0 repositories listed
-
Full-text Error Correction for Chinese Speech Recognition with Large Language Model12 Sep 2024 0 repositories listed
-
Enhancing CTC-Based Visual Speech Recognition11 Sep 2024 0 repositories listed
-
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition10 Sep 2024 0 repositories listed
-
Keyword-Aware ASR Error Augmentation for Robust Dialogue State Tracking10 Sep 2024 0 repositories listed
-
An investigation of modularity for noise robustness in conformer-based ASR9 Sep 2024 0 repositories listed
-
Evaluation of real-time transcriptions using end-to-end ASR models9 Sep 2024 0 repositories listed
-
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge9 Sep 2024 0 repositories listed
-
Retrieval Augmented Correction of Named Entity Speech Recognition Errors9 Sep 2024 0 repositories listed
-
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection8 Sep 2024 0 repositories listed
-
Probing self-attention in self-supervised speech models for cross-linguistic differences4 Sep 2024 0 repositories listed
-
Quantification of stylistic differences in human- and ASR-produced transcripts of African American English4 Sep 2024 0 repositories listed
-
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations4 Sep 2024 0 repositories listed
-
Reassessing Noise Augmentation Methods in the Context of Adversarial Speech3 Sep 2024 0 repositories listed
-
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR3 Sep 2024 0 repositories listed
-
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka3 Sep 2024 0 repositories listed
-
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR2 Sep 2024 0 repositories listed
-
Comparing Discrete and Continuous Space LLMs for Speech Recognition1 Sep 2024 0 repositories listed
-
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition1 Sep 2024 0 repositories listed
-
Advancing Multi-talker ASR Performance with Large Language Models30 Aug 2024 0 repositories listed
-
Speaker Tagging Correction With Non-Autoregressive Language Models30 Aug 2024 0 repositories listed
-
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction29 Aug 2024 0 repositories listed
-
Automatic recognition and detection of aphasic natural speech26 Aug 2024 0 repositories listed
-
MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues26 Aug 2024 0 repositories listed
-
Focused Discriminative Training For Streaming CTC-Trained Automatic Speech Recognition Models23 Aug 2024 0 repositories listed
-
The State of Commercial Automatic French Legal Speech Recognition Systems and their Impact on Court Reporters et al21 Aug 2024 0 repositories listed
-
Parameter-Efficient Transfer Learning under Federated Learning for Automatic Speech Recognition19 Aug 2024 0 repositories listed
-
Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts19 Aug 2024 0 repositories listed
-
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words15 Aug 2024 0 repositories listed
-
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation13 Aug 2024 0 repositories listed
-
Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance12 Aug 2024 0 repositories listed
-
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning12 Aug 2024 0 repositories listed
-
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing11 Aug 2024 0 repositories listed
-
7 Aug 2024 0 repositories listed
-
Self-Supervised Learning for Multi-Channel Neural Transducer6 Aug 2024 0 repositories listed
-
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval6 Aug 2024 0 repositories listed
-
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion5 Aug 2024 0 repositories listed
-
Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation1 Aug 2024 0 repositories listed
-
Towards interfacing large language models with ASR systems using confidence measures and prompting31 Jul 2024 0 repositories listed
-
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition31 Jul 2024 0 repositories listed
-
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures25 Jul 2024 0 repositories listed
-
Reexamining Racial Disparities in Automatic Speech Recognition Performance: The Role of Confounding by Provenance19 Jul 2024 0 repositories listed
-
Handling Numeric Expressions in Automatic Speech Recognition18 Jul 2024 0 repositories listed
-
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR18 Jul 2024 0 repositories listed
-
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training18 Jul 2024 0 repositories listed
-
Robust ASR Error Correction with Conservative Data Filtering18 Jul 2024 0 repositories listed
-
Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models16 Jul 2024 0 repositories listed