Browse State-of-the-Art › Automatic Speech Recognition (ASR) › Papers, page 8
Automatic Speech Recognition (ASR)
Papers archive 2025-07-28
archive papers tagged: 3,012 · with a code link: 622 · where Syntology ran a sample: 77 (64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (77 of 3,012 tagged: 64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument)
Page 8 of 31: papers 701 to 800 of 3,012, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
The Impact of Code-switched Synthetic Data Quality is Task Dependent: Insights from MT and ASR30 Mar 2025 0 repositories listed
-
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications26 Mar 2025 0 repositories listed
-
Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization25 Mar 2025 0 repositories listed
-
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language24 Mar 2025 0 repositories listed
-
Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication21 Mar 2025 0 repositories listed
-
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces19 Mar 2025 0 repositories listed
-
Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment12 Mar 2025 0 repositories listed
-
ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization12 Mar 2025 0 repositories listed
-
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR11 Mar 2025 0 repositories listed
-
Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling10 Mar 2025 0 repositories listed
-
Building English ASR model with regional language support10 Mar 2025 0 repositories listed
-
From Voice to Safety: Language AI Powered Pilot-ATC Communication Understanding for Airport Surface Movement Collision Risk Assessment6 Mar 2025 0 repositories listed
-
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations5 Mar 2025 0 repositories listed
-
Direct Speech to Speech Translation: A Review3 Mar 2025 0 repositories listed
-
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis3 Mar 2025 0 repositories listed
-
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems2 Mar 2025 0 repositories listed
-
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications27 Feb 2025 0 repositories listed
-
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition26 Feb 2025 0 repositories listed
-
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision26 Feb 2025 0 repositories listed
-
Exploring Gender Disparities in Automatic Speech Recognition Technology25 Feb 2025 0 repositories listed
-
Understanding Zero-shot Rare Word Recognition Improvements Through LLM Integration22 Feb 2025 0 repositories listed
-
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders21 Feb 2025 0 repositories listed
-
The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages21 Feb 2025 0 repositories listed
-
Adopting Whisper for Confidence Estimation19 Feb 2025 0 repositories listed
-
Benchmarking Automatic Speech Recognition coupled LLM Modules for Medical Diagnostics18 Feb 2025 0 repositories listed
-
Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders18 Feb 2025 0 repositories listed
-
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models18 Feb 2025 0 repositories listed
-
MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems14 Feb 2025 0 repositories listed
-
Causal Analysis of ASR Errors for Children: Quantifying the Impact of Physiological, Cognitive, and Extrinsic Factors12 Feb 2025 0 repositories listed
-
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance7 Feb 2025 0 repositories listed
-
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond6 Feb 2025 0 repositories listed
-
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport3 Feb 2025 0 repositories listed
-
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition3 Feb 2025 0 repositories listed
-
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition1 Feb 2025 0 repositories listed
-
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation1 Feb 2025 0 repositories listed
-
Language Bias in Self-Supervised Learning For Automatic Speech Recognition31 Jan 2025 0 repositories listed
-
SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions31 Jan 2025 0 repositories listed
-
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition29 Jan 2025 0 repositories listed
-
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation26 Jan 2025 0 repositories listed
-
The Multicultural Medical Assistant: Can LLMs Improve Medical ASR Errors Across Borders?25 Jan 2025 0 repositories listed
-
Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing23 Jan 2025 0 repositories listed
-
22 Jan 2025 0 repositories listed
-
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio20 Jan 2025 0 repositories listed
-
A Benchmark of French ASR Systems Based on Error Severity18 Jan 2025 0 repositories listed
-
GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems18 Jan 2025 0 repositories listed
-
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR17 Jan 2025 0 repositories listed
-
Adapting Whisper for Regional Dialects: Enhancing Public Services for Vulnerable Populations in the United Kingdom15 Jan 2025 0 repositories listed
-
persoDA: Personalized Data Augmentation for Personalized ASR15 Jan 2025 0 repositories listed
-
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives11 Jan 2025 0 repositories listed
-
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition10 Jan 2025 0 repositories listed
-
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI10 Jan 2025 0 repositories listed
-
Universal-2-TF: Robust All-Neural Text Formatting for ASR10 Jan 2025 0 repositories listed
-
6 Jan 2025 0 repositories listed
-
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer3 Jan 2025 0 repositories listed
-
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale1 Jan 2025 0 repositories listed
-
Zero-resource Speech Translation and Recognition with LLMs24 Dec 2024 0 repositories listed
-
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition23 Dec 2024 0 repositories listed
-
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding21 Dec 2024 0 repositories listed
-
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling21 Dec 2024 0 repositories listed
-
Speech Retrieval-Augmented Generation without Automatic Speech Recognition21 Dec 2024 0 repositories listed
-
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition21 Dec 2024 0 repositories listed
-
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch20 Dec 2024 0 repositories listed
-
LAMA-UT: Language Agnostic Multilingual ASR through Orthography Unification and Language-Specific Transliteration19 Dec 2024 0 repositories listed
-
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition19 Dec 2024 0 repositories listed
-
Speak & Improve Challenge 2025: Tasks and Baseline Systems16 Dec 2024 0 repositories listed
-
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback16 Dec 2024 0 repositories listed
-
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation11 Dec 2024 0 repositories listed
-
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning9 Dec 2024 0 repositories listed
-
Harnessing Transfer Learning from Swahili: Advancing Solutions for Comorian Dialects9 Dec 2024 0 repositories listed
-
Leveraging Prompt Learning and Pause Encoding for Alzheimer's Disease Detection9 Dec 2024 0 repositories listed
-
Not All Errors Are Equal: Investigation of Speech Recognition Errors in Alzheimer's Disease Detection9 Dec 2024 0 repositories listed
-
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding5 Dec 2024 0 repositories listed
-
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction4 Dec 2024 0 repositories listed
-
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario1 Dec 2024 0 repositories listed
-
Late fusion ensembles for speech recognition on diverse input audio representations1 Dec 2024 0 repositories listed
-
Empowering the Deaf and Hard of Hearing Community: Enhancing Video Captions Using Large Language Models30 Nov 2024 0 repositories listed
-
Aligning Pre-trained Models for Spoken Language Translation27 Nov 2024 0 repositories listed
-
AMPS: ASR with Multimodal Paraphrase Supervision27 Nov 2024 0 repositories listed
-
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory27 Nov 2024 0 repositories listed
-
How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario27 Nov 2024 0 repositories listed
-
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation26 Nov 2024 0 repositories listed
-
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection26 Nov 2024 0 repositories listed
-
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition26 Nov 2024 0 repositories listed
-
24 Nov 2024 0 repositories listed
-
Transforming NLU with Babylon: A Case Study in Development of Real-time, Edge-Efficient, Multi-Intent Translation System for Automated Drive-Thru Ordering22 Nov 2024 0 repositories listed
-
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge21 Nov 2024 0 repositories listed
-
CAFE A Novel Code switching Dataset for Algerian Dialect French and English20 Nov 2024 0 repositories listed
-
From Statistical Methods to Pre-Trained Models; A Survey on Automatic Speech Recognition for Resource Scarce Urdu Language20 Nov 2024 0 repositories listed
-
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM20 Nov 2024 0 repositories listed
-
Whisper Finetuning on Nepali Language19 Nov 2024 0 repositories listed
-
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data14 Nov 2024 0 repositories listed
-
Transferable Adversarial Attacks against ASR14 Nov 2024 0 repositories listed
-
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions11 Nov 2024 0 repositories listed
-
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages7 Nov 2024 0 repositories listed
-
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO1 Nov 2024 0 repositories listed
-
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising30 Oct 2024 0 repositories listed
-
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs27 Oct 2024 0 repositories listed
-
A Survey on Speech Large Language Models24 Oct 2024 0 repositories listed
-
Evaluating and Improving Automatic Speech Recognition Systems for Korean Meteorological Experts24 Oct 2024 0 repositories listed
-
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams23 Oct 2024 0 repositories listed