Browse State-of-the-Art › speech-recognition › Papers, page 15
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 15 of 58: papers 1,401 to 1,500 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models16 May 2025 0 repositories listed
-
Inclusivity of AI Speech in Healthcare: A Decade Look Back15 May 2025 0 repositories listed
-
Quantized Approximate Signal Processing (QASP): Towards Homomorphic Encryption for audio15 May 2025 0 repositories listed
-
Full simulation on the dynamics of auditory synaptic fusion: Strong clustering of calcium channel might be the origin of the coherent release in the auditory hair cells12 May 2025 0 repositories listed
-
Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients9 May 2025 0 repositories listed
-
Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations8 May 2025 0 repositories listed
-
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement7 May 2025 0 repositories listed
-
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer7 May 2025 0 repositories listed
-
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech6 May 2025 0 repositories listed
-
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation6 May 2025 0 repositories listed
-
A Synergistic Framework of Nonlinear Acoustic Computing and Reinforcement Learning for Real-World Human-Robot Interaction4 May 2025 0 repositories listed
-
Transfer Learning-Based Deep Residual Learning for Speech Recognition in Clean and Noisy Environments2 May 2025 0 repositories listed
-
Retrieval-Enhanced Few-Shot Prompting for Speech Event Extraction30 Apr 2025 0 repositories listed
-
Development and evaluation of a deep learning algorithm for German word recognition from lip movements22 Apr 2025 0 repositories listed
-
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides21 Apr 2025 0 repositories listed
-
StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models21 Apr 2025 0 repositories listed
-
Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope17 Apr 2025 0 repositories listed
-
Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning16 Apr 2025 0 repositories listed
-
Spatial Audio Processing with Large Language Model on Wearable Devices11 Apr 2025 0 repositories listed
-
Summarizing Speech: A Comprehensive Survey10 Apr 2025 0 repositories listed
-
Visual-Aware Speech Recognition for Noisy Scenarios9 Apr 2025 0 repositories listed
-
A Human Digital Twin Architecture for Knowledge-based Interactions and Context-Aware Conversations4 Apr 2025 0 repositories listed
-
Edge Intelligence for Wildlife Conservation: Real-Time Hornbill Call Classification Using TinyML3 Apr 2025 0 repositories listed
-
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect3 Apr 2025 0 repositories listed
-
Chain of Correction for Full-text Speech Recognition with Large Language Models2 Apr 2025 0 repositories listed
-
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models30 Mar 2025 0 repositories listed
-
The Impact of Code-switched Synthetic Data Quality is Task Dependent: Insights from MT and ASR30 Mar 2025 0 repositories listed
-
A 71.2-μW Speech Recognition Accelerator with Recurrent Spiking Neural Network27 Mar 2025 0 repositories listed
-
VALLR: Visual ASR Language Model for Lip Reading27 Mar 2025 0 repositories listed
-
Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance26 Mar 2025 0 repositories listed
-
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications26 Mar 2025 0 repositories listed
-
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit26 Mar 2025 0 repositories listed
-
Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization25 Mar 2025 0 repositories listed
-
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy25 Mar 2025 0 repositories listed
-
Coverage-Guaranteed Speech Emotion Recognition via Calibrated Uncertainty-Adaptive Prediction Sets24 Mar 2025 0 repositories listed
-
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language24 Mar 2025 0 repositories listed
-
From S4 to Mamba: A Comprehensive Survey on Structured State Space Models22 Mar 2025 0 repositories listed
-
Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication21 Mar 2025 0 repositories listed
-
SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors20 Mar 2025 0 repositories listed
-
A Comprehensive Survey on Architectural Advances in Deep CNNs: Challenges, Applications, and Emerging Research Directions19 Mar 2025 0 repositories listed
-
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces19 Mar 2025 0 repositories listed
-
Halving transcription time: A fast, user-friendly and GDPR-compliant workflow to create AI-assisted transcripts for content analysis17 Mar 2025 0 repositories listed
-
Enhancing Aviation Communication Transcription: Fine-Tuning Distil-Whisper with LoRA13 Mar 2025 0 repositories listed
-
Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment12 Mar 2025 0 repositories listed
-
Proceedings of the ISCA/ITG Workshop on Diversity in Large Speech and Language Models12 Mar 2025 0 repositories listed
-
ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization12 Mar 2025 0 repositories listed
-
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR11 Mar 2025 0 repositories listed
-
Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling10 Mar 2025 0 repositories listed
-
Building English ASR model with regional language support10 Mar 2025 0 repositories listed
-
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs9 Mar 2025 0 repositories listed
-
A Causal Inference Approach for Quantifying Research Impact7 Mar 2025 0 repositories listed
-
From Voice to Safety: Language AI Powered Pilot-ATC Communication Understanding for Airport Surface Movement Collision Risk Assessment6 Mar 2025 0 repositories listed
-
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning6 Mar 2025 0 repositories listed
-
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations5 Mar 2025 0 repositories listed
-
CORDIC Is All You Need4 Mar 2025 0 repositories listed
-
Direct Speech to Speech Translation: A Review3 Mar 2025 0 repositories listed
-
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis3 Mar 2025 0 repositories listed
-
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation2 Mar 2025 0 repositories listed
-
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems2 Mar 2025 0 repositories listed
-
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications27 Feb 2025 0 repositories listed
-
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition26 Feb 2025 0 repositories listed
-
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision26 Feb 2025 0 repositories listed
-
Exploring Gender Disparities in Automatic Speech Recognition Technology25 Feb 2025 0 repositories listed
-
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM24 Feb 2025 0 repositories listed
-
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation24 Feb 2025 0 repositories listed
-
Understanding Zero-shot Rare Word Recognition Improvements Through LLM Integration22 Feb 2025 0 repositories listed
-
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders21 Feb 2025 0 repositories listed
-
Retrieval-Augmented Speech Recognition Approach for Domain Challenges21 Feb 2025 0 repositories listed
-
The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages21 Feb 2025 0 repositories listed
-
Moshi Moshi? A Model Selection Hijacking Adversarial Attack20 Feb 2025 0 repositories listed
-
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models20 Feb 2025 0 repositories listed
-
Adopting Whisper for Confidence Estimation19 Feb 2025 0 repositories listed
-
Benchmarking Automatic Speech Recognition coupled LLM Modules for Medical Diagnostics18 Feb 2025 0 repositories listed
-
Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders18 Feb 2025 0 repositories listed
-
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models18 Feb 2025 0 repositories listed
-
Neuro-oscillatory models of cortical speech processing18 Feb 2025 0 repositories listed
-
On the Robust Approximation of ASR Metrics18 Feb 2025 0 repositories listed
-
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization18 Feb 2025 0 repositories listed
-
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing17 Feb 2025 0 repositories listed
-
A Preliminary Exploration with GPT-4o Voice Mode14 Feb 2025 0 repositories listed
-
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge14 Feb 2025 0 repositories listed
-
MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems14 Feb 2025 0 repositories listed
-
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models14 Feb 2025 0 repositories listed
-
Shortcut Learning Susceptibility in Vision Classifiers13 Feb 2025 0 repositories listed
-
Causal Analysis of ASR Errors for Children: Quantifying the Impact of Physiological, Cognitive, and Extrinsic Factors12 Feb 2025 0 repositories listed
-
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition11 Feb 2025 0 repositories listed
-
Speech to Speech Translation with Translatotron: A State of the Art Review9 Feb 2025 0 repositories listed
-
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance7 Feb 2025 0 repositories listed
-
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance7 Feb 2025 0 repositories listed
-
Lightweight Operations for Visual Speech Recognition7 Feb 2025 0 repositories listed
-
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond6 Feb 2025 0 repositories listed
-
Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers6 Feb 2025 0 repositories listed
-
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport3 Feb 2025 0 repositories listed
-
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models3 Feb 2025 0 repositories listed
-
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition3 Feb 2025 0 repositories listed
-
Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition3 Feb 2025 0 repositories listed
-
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition1 Feb 2025 0 repositories listed
-
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation1 Feb 2025 0 repositories listed
-
Language Bias in Self-Supervised Learning For Automatic Speech Recognition31 Jan 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.