Browse State-of-the-Art › speech-recognition › Papers, page 16
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 16 of 58: papers 1,501 to 1,600 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions31 Jan 2025 0 repositories listed
-
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition29 Jan 2025 0 repositories listed
-
Privacy-Preserving Edge Speech Understanding with Tiny Foundation Models29 Jan 2025 0 repositories listed
-
SCDiar: a streaming diarization system based on speaker change detection and speech recognition28 Jan 2025 0 repositories listed
-
Classification Error Bound for Low Bayes Error Conditions in Machine Learning27 Jan 2025 0 repositories listed
-
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario26 Jan 2025 0 repositories listed
-
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation26 Jan 2025 0 repositories listed
-
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition25 Jan 2025 0 repositories listed
-
The Multicultural Medical Assistant: Can LLMs Improve Medical ASR Errors Across Borders?25 Jan 2025 0 repositories listed
-
LoCoML: A Framework for Real-World ML Inference Pipelines24 Jan 2025 0 repositories listed
-
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition23 Jan 2025 0 repositories listed
-
Integrating Persian Lip Reading in Surena-V Humanoid Robot for Human-Robot Interaction23 Jan 2025 0 repositories listed
-
Learning-based A Posteriori Speech Presence Probability Estimation and Applications23 Jan 2025 0 repositories listed
-
Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing23 Jan 2025 0 repositories listed
-
Development of an Inclusive Educational Platform Using Open Technologies and Machine Learning: A Case Study on Accessibility Enhancement22 Jan 2025 0 repositories listed
-
22 Jan 2025 0 repositories listed
-
A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data21 Jan 2025 0 repositories listed
-
Generative AI and Large Language Models in Language Preservation: Opportunities and Challenges20 Jan 2025 0 repositories listed
-
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio20 Jan 2025 0 repositories listed
-
Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets19 Jan 2025 0 repositories listed
-
A Benchmark of French ASR Systems Based on Error Severity18 Jan 2025 0 repositories listed
-
GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems18 Jan 2025 0 repositories listed
-
Automatic Speech Recognition for Sanskrit with Transfer Learning17 Jan 2025 0 repositories listed
-
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR17 Jan 2025 0 repositories listed
-
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition16 Jan 2025 0 repositories listed
-
A Non-autoregressive Model for Joint STT and TTS15 Jan 2025 0 repositories listed
-
Adapting Whisper for Regional Dialects: Enhancing Public Services for Vulnerable Populations in the United Kingdom15 Jan 2025 0 repositories listed
-
persoDA: Personalized Data Augmentation for Personalized ASR15 Jan 2025 0 repositories listed
-
Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications14 Jan 2025 0 repositories listed
-
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model13 Jan 2025 0 repositories listed
-
A Survey on Spoken Italian Datasets and Corpora11 Jan 2025 0 repositories listed
-
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives11 Jan 2025 0 repositories listed
-
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition10 Jan 2025 0 repositories listed
-
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI10 Jan 2025 0 repositories listed
-
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer10 Jan 2025 0 repositories listed
-
Universal-2-TF: Robust All-Neural Text Formatting for ASR10 Jan 2025 0 repositories listed
-
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition8 Jan 2025 0 repositories listed
-
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages8 Jan 2025 0 repositories listed
-
Deep Learning for Pathological Speech: A Survey7 Jan 2025 0 repositories listed
-
Towards a Generalizable Speech Marker for Parkinson's Disease Diagnosis7 Jan 2025 0 repositories listed
-
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection7 Jan 2025 0 repositories listed
-
6 Jan 2025 0 repositories listed
-
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer3 Jan 2025 0 repositories listed
-
Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing1 Jan 2025 0 repositories listed
-
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition1 Jan 2025 0 repositories listed
-
Incremental Dialogue Management: Survey, Discussion, and Implications for HRI1 Jan 2025 0 repositories listed
-
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale1 Jan 2025 0 repositories listed
-
Fotheidil: an Automatic Transcription System for the Irish Language31 Dec 2024 0 repositories listed
-
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages31 Dec 2024 0 repositories listed
-
Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization27 Dec 2024 0 repositories listed
-
Towards a Single ASR Model That Generalizes to Disordered Speech26 Dec 2024 0 repositories listed
-
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning25 Dec 2024 0 repositories listed
-
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition25 Dec 2024 0 repositories listed
-
Zero-resource Speech Translation and Recognition with LLMs24 Dec 2024 0 repositories listed
-
Deep Learning in Proteomics Informatics: Applications, Challenges, and Future Directions23 Dec 2024 0 repositories listed
-
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution23 Dec 2024 0 repositories listed
-
Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning23 Dec 2024 0 repositories listed
-
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition23 Dec 2024 0 repositories listed
-
Uncovering the Visual Contribution in Audio-Visual Speech Recognition22 Dec 2024 0 repositories listed
-
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding21 Dec 2024 0 repositories listed
-
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling21 Dec 2024 0 repositories listed
-
Speech Retrieval-Augmented Generation without Automatic Speech Recognition21 Dec 2024 0 repositories listed
-
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition21 Dec 2024 0 repositories listed
-
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch20 Dec 2024 0 repositories listed
-
LAMA-UT: Language Agnostic Multilingual ASR through Orthography Unification and Language-Specific Transliteration19 Dec 2024 0 repositories listed
-
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition19 Dec 2024 0 repositories listed
-
Speak & Improve Challenge 2025: Tasks and Baseline Systems16 Dec 2024 0 repositories listed
-
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback16 Dec 2024 0 repositories listed
-
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond16 Dec 2024 0 repositories listed
-
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition15 Dec 2024 0 repositories listed
-
Robust Persian Digit Recognition in Noisy Environments Using Hybrid CNN-BiGRU Model14 Dec 2024 0 repositories listed
-
MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models13 Dec 2024 0 repositories listed
-
Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition11 Dec 2024 0 repositories listed
-
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation11 Dec 2024 0 repositories listed
-
Style-agnostic evaluation of ASR using multiple reference transcripts10 Dec 2024 0 repositories listed
-
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning9 Dec 2024 0 repositories listed
-
Ensemble Machine Learning Model for Inner Speech Recognition: A Subject-Specific Investigation9 Dec 2024 0 repositories listed
-
Harnessing Transfer Learning from Swahili: Advancing Solutions for Comorian Dialects9 Dec 2024 0 repositories listed
-
Leveraging Prompt Learning and Pause Encoding for Alzheimer's Disease Detection9 Dec 2024 0 repositories listed
-
Not All Errors Are Equal: Investigation of Speech Recognition Errors in Alzheimer's Disease Detection9 Dec 2024 0 repositories listed
-
Adaptive Dropout for Pruning Conformers6 Dec 2024 0 repositories listed
-
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding5 Dec 2024 0 repositories listed
-
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech5 Dec 2024 0 repositories listed
-
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction4 Dec 2024 0 repositories listed
-
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario1 Dec 2024 0 repositories listed
-
Late fusion ensembles for speech recognition on diverse input audio representations1 Dec 2024 0 repositories listed
-
Empowering the Deaf and Hard of Hearing Community: Enhancing Video Captions Using Large Language Models30 Nov 2024 0 repositories listed
-
28 Nov 2024 0 repositories listed
-
Aligning Pre-trained Models for Spoken Language Translation27 Nov 2024 0 repositories listed
-
AMPS: ASR with Multimodal Paraphrase Supervision27 Nov 2024 0 repositories listed
-
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory27 Nov 2024 0 repositories listed
-
How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario27 Nov 2024 0 repositories listed
-
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models27 Nov 2024 0 repositories listed
-
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation27 Nov 2024 0 repositories listed
-
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation26 Nov 2024 0 repositories listed
-
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection26 Nov 2024 0 repositories listed
-
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition26 Nov 2024 0 repositories listed
-
24 Nov 2024 0 repositories listed
-
Transforming NLU with Babylon: A Case Study in Development of Real-time, Edge-Efficient, Multi-Intent Translation System for Automated Drive-Thru Ordering22 Nov 2024 0 repositories listed
-
TSkips: Efficiency Through Explicit Temporal Delay Connections in Spiking Neural Networks22 Nov 2024 0 repositories listed