Browse State-of-the-Art › Automatic Speech Recognition › Papers, page 12
Automatic Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 3,174 · with a code link: 677 · where Syntology ran a sample: 79 (62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (79 of 3,174 tagged: 62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 12 of 32: papers 1,101 to 1,200 of 3,174, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores6 Jun 2024 0 repositories listed
-
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders5 Jun 2024 0 repositories listed
-
Enhancing CTC-based speech recognition with diverse modeling units5 Jun 2024 0 repositories listed
-
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition5 Jun 2024 0 repositories listed
-
Text Injection for Neural Contextual Biasing5 Jun 2024 0 repositories listed
-
Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping4 Jun 2024 0 repositories listed
-
Keyword-Guided Adaptation of Automatic Speech Recognition4 Jun 2024 0 repositories listed
-
Enabling ASR for Low-Resource Languages: A Comprehensive Dataset Creation Approach3 Jun 2024 0 repositories listed
-
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning1 Jun 2024 0 repositories listed
-
Zipper: A Multi-Tower Decoder Architecture for Fusing Modalities29 May 2024 0 repositories listed
-
Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation28 May 2024 0 repositories listed
-
Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition24 May 2024 0 repositories listed
-
Contextualized Automatic Speech Recognition with Dynamic Vocabulary22 May 2024 0 repositories listed
-
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation22 May 2024 0 repositories listed
-
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish22 May 2024 0 repositories listed
-
FairLENS: Assessing Fairness in Law Enforcement Speech Recognition21 May 2024 0 repositories listed
-
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models16 May 2024 0 repositories listed
-
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings15 May 2024 0 repositories listed
-
Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer15 May 2024 0 repositories listed
-
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants14 May 2024 0 repositories listed
-
SpeechVerse: A Large-scale Generalizable Audio Language Model14 May 2024 0 repositories listed
-
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech10 May 2024 0 repositories listed
-
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition6 May 2024 0 repositories listed
-
Combining X-Vectors and Bayesian Batch Active Learning: Two-Stage Active Learning Pipeline for Speech Recognition3 May 2024 0 repositories listed
-
Efficient Compression of Multitask Multilingual Speech Models2 May 2024 0 repositories listed
-
Improving Membership Inference in ASR Model Auditing with Perturbed Loss Features2 May 2024 0 repositories listed
-
Sequence-to-sequence models in peer-to-peer learning: A practical application2 May 2024 0 repositories listed
-
Does Whisper understand Swiss German? An automatic, qualitative, and human evaluation30 Apr 2024 0 repositories listed
-
Automatic Speech Recognition System-Independent Word Error Rate Estimation25 Apr 2024 0 repositories listed
-
Developing Acoustic Models for Automatic Speech Recognition in Swedish25 Apr 2024 0 repositories listed
-
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF25 Apr 2024 0 repositories listed
-
Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices24 Apr 2024 0 repositories listed
-
Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm23 Apr 2024 0 repositories listed
-
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance23 Apr 2024 0 repositories listed
-
Efficient infusion of self-supervised representations in Automatic Speech Recognition19 Apr 2024 0 repositories listed
-
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech18 Apr 2024 0 repositories listed
-
Anatomy of Industrial Scale Multilingual ASR15 Apr 2024 0 repositories listed
-
Resilience of Large Language Models for Noisy Instructions15 Apr 2024 0 repositories listed
-
12 Apr 2024 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task12 Apr 2024 0 repositories listed
-
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution11 Apr 2024 0 repositories listed
-
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping10 Apr 2024 0 repositories listed
-
The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge9 Apr 2024 0 repositories listed
-
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition4 Apr 2024 0 repositories listed
-
Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian3 Apr 2024 0 repositories listed
-
Noise Masking Attacks and Defenses for Pretrained Speech Models2 Apr 2024 0 repositories listed
-
Transfer Learning from Whisper for Microscopic Intelligibility Prediction2 Apr 2024 0 repositories listed
-
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models31 Mar 2024 0 repositories listed
-
LV-CTC: Non-autoregressive ASR with CTC and latent variable models28 Mar 2024 0 repositories listed
-
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition28 Mar 2024 0 repositories listed
-
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus27 Mar 2024 0 repositories listed
-
DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech Recognition26 Mar 2024 0 repositories listed
-
Extracting Biomedical Entities from Noisy Audio Transcripts26 Mar 2024 0 repositories listed
-
Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models25 Mar 2024 0 repositories listed
-
A Multimodal Approach to Device-Directed Speech Detection with Large Language Models21 Mar 2024 0 repositories listed
-
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech20 Mar 2024 0 repositories listed
-
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning20 Mar 2024 0 repositories listed
-
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition18 Mar 2024 0 repositories listed
-
Artificial Intelligence for Cochlear Implants: Review of Strategies, Challenges, and Perspectives17 Mar 2024 0 repositories listed
-
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children13 Mar 2024 0 repositories listed
-
Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition13 Mar 2024 0 repositories listed
-
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language12 Mar 2024 0 repositories listed
-
The evaluation of a code-switched Sepedi-English automatic speech recognition system11 Mar 2024 0 repositories listed
-
Aligning Speech to Languages to Enhance Code-switching Speech Recognition9 Mar 2024 0 repositories listed
-
A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain7 Mar 2024 0 repositories listed
-
Classist Tools: Social Class Correlates with Performance in NLP7 Mar 2024 0 repositories listed
-
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition4 Mar 2024 0 repositories listed
-
What has LeBenchmark Learnt about French Syntax?4 Mar 2024 0 repositories listed
-
A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement3 Mar 2024 0 repositories listed
-
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey2 Mar 2024 0 repositories listed
-
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview1 Mar 2024 0 repositories listed
-
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition29 Feb 2024 0 repositories listed
-
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems29 Feb 2024 0 repositories listed
-
Exploration of Adapter for Noise Robust Automatic Speech Recognition28 Feb 2024 0 repositories listed
-
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement27 Feb 2024 0 repositories listed
-
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models27 Feb 2024 0 repositories listed
-
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR21 Feb 2024 0 repositories listed
-
Ain't Misbehavin' -- Using LLMs to Generate Expressive Robot Behavior in Conversations with the Tabletop Robot Haru18 Feb 2024 0 repositories listed
-
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models14 Feb 2024 0 repositories listed
-
The Balancing Act: Unmasking and Alleviating ASR Biases in Portuguese12 Feb 2024 0 repositories listed
-
The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models12 Feb 2024 0 repositories listed
-
Self-consistent context aware conformer transducer for speech recognition9 Feb 2024 0 repositories listed
-
Progressive unsupervised domain adaptation for ASR using ensemble models and multi-stage training7 Feb 2024 0 repositories listed
-
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems5 Feb 2024 0 repositories listed
-
Resolving Transcription Ambiguity in Spanish: A Hybrid Acoustic-Lexical System for Punctuation Restoration5 Feb 2024 0 repositories listed
-
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens3 Feb 2024 0 repositories listed
-
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents2 Feb 2024 0 repositories listed
-
Digits micro-model for accurate and secure transactions2 Feb 2024 0 repositories listed
-
Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges2 Feb 2024 0 repositories listed
-
Byte Pair Encoding Is All You Need For Automatic Bengali Speech Recognition28 Jan 2024 0 repositories listed
-
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline26 Jan 2024 0 repositories listed
-
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction24 Jan 2024 0 repositories listed
-
Locality enhanced dynamic biasing and sampling strategies for contextual ASR23 Jan 2024 0 repositories listed
-
Consistency Based Unsupervised Self-training For ASR Personalisation22 Jan 2024 0 repositories listed
-
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers22 Jan 2024 0 repositories listed
-
Using Large Language Model for End-to-End Chinese ASR and NER21 Jan 2024 0 repositories listed
-
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search19 Jan 2024 0 repositories listed
-
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition18 Jan 2024 0 repositories listed
-
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks18 Jan 2024 0 repositories listed
-
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition18 Jan 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.