Browse State-of-the-Art › Automatic Speech Recognition (ASR) › Papers, page 10
Automatic Speech Recognition (ASR)
Papers archive 2025-07-28
archive papers tagged: 3,012 · with a code link: 622 · where Syntology ran a sample: 77 (64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (77 of 3,012 tagged: 64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument)
Page 10 of 31: papers 901 to 1,000 of 3,012, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data15 Jul 2024 0 repositories listed
-
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation14 Jul 2024 0 repositories listed
-
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing10 Jul 2024 0 repositories listed
-
Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation8 Jul 2024 0 repositories listed
-
LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech5 Jul 2024 0 repositories listed
-
Romanization Encoding For Multilingual ASR5 Jul 2024 0 repositories listed
-
5 Jul 2024 0 repositories listed
-
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter5 Jul 2024 0 repositories listed
-
Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models5 Jul 2024 0 repositories listed
-
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis4 Jul 2024 0 repositories listed
-
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations3 Jul 2024 0 repositories listed
-
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition29 Jun 2024 0 repositories listed
-
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over27 Jun 2024 0 repositories listed
-
Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment27 Jun 2024 0 repositories listed
-
Automatic Speech Recognition for Hindi26 Jun 2024 0 repositories listed
-
Dynamic Data Pruning for Automatic Speech Recognition26 Jun 2024 0 repositories listed
-
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research26 Jun 2024 0 repositories listed
-
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR26 Jun 2024 0 repositories listed
-
Sequential Editing for Lifelong Training of Speech Recognition Models25 Jun 2024 0 repositories listed
-
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 202424 Jun 2024 0 repositories listed
-
Decoder-only Architecture for Streaming End-to-end Speech Recognition23 Jun 2024 0 repositories listed
-
Perception of Phonological Assimilation by Neural Speech Recognition Models21 Jun 2024 0 repositories listed
-
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices21 Jun 2024 0 repositories listed
-
Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control19 Jun 2024 0 repositories listed
-
ManWav: The First Manchu ASR Model19 Jun 2024 0 repositories listed
-
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model18 Jun 2024 0 repositories listed
-
Performant ASR Models for Medical Entities in Accented Speech18 Jun 2024 0 repositories listed
-
Transcribe, Align and Segment: Creating speech datasets for low-resource languages18 Jun 2024 0 repositories listed
-
Automatic Speech Recognition for Biomedical Data in Bengali Language16 Jun 2024 0 repositories listed
-
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving16 Jun 2024 0 repositories listed
-
An efficient text augmentation approach for contextualized Mandarin speech recognition14 Jun 2024 0 repositories listed
-
Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation14 Jun 2024 0 repositories listed
-
Optimizing Byte-level Representation for End-to-end ASR14 Jun 2024 0 repositories listed
-
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR14 Jun 2024 0 repositories listed
-
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment13 Jun 2024 0 repositories listed
-
The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments13 Jun 2024 0 repositories listed
-
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition13 Jun 2024 0 repositories listed
-
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data12 Jun 2024 0 repositories listed
-
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion12 Jun 2024 0 repositories listed
-
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets12 Jun 2024 0 repositories listed
-
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding12 Jun 2024 0 repositories listed
-
Transformer-based Model for ASR N-Best Rescoring and Rewriting12 Jun 2024 0 repositories listed
-
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection11 Jun 2024 0 repositories listed
-
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter11 Jun 2024 0 repositories listed
-
Reading Miscue Detection in Primary School through Automatic Speech Recognition11 Jun 2024 0 repositories listed
-
ASTRA: Aligning Speech and Text Representations for Asr without Sampling10 Jun 2024 0 repositories listed
-
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations9 Jun 2024 0 repositories listed
-
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR7 Jun 2024 0 repositories listed
-
Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis7 Jun 2024 0 repositories listed
-
Flexible Multichannel Speech Enhancement for Noise-Robust Frontend6 Jun 2024 0 repositories listed
-
Hypernetworks for Personalizing ASR to Atypical Speech6 Jun 2024 0 repositories listed
-
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores6 Jun 2024 0 repositories listed
-
Enhancing CTC-based speech recognition with diverse modeling units5 Jun 2024 0 repositories listed
-
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition5 Jun 2024 0 repositories listed
-
Text Injection for Neural Contextual Biasing5 Jun 2024 0 repositories listed
-
Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping4 Jun 2024 0 repositories listed
-
Keyword-Guided Adaptation of Automatic Speech Recognition4 Jun 2024 0 repositories listed
-
Enabling ASR for Low-Resource Languages: A Comprehensive Dataset Creation Approach3 Jun 2024 0 repositories listed
-
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning1 Jun 2024 0 repositories listed
-
Zipper: A Multi-Tower Decoder Architecture for Fusing Modalities29 May 2024 0 repositories listed
-
Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation28 May 2024 0 repositories listed
-
Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition24 May 2024 0 repositories listed
-
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation22 May 2024 0 repositories listed
-
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish22 May 2024 0 repositories listed
-
FairLENS: Assessing Fairness in Law Enforcement Speech Recognition21 May 2024 0 repositories listed
-
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models16 May 2024 0 repositories listed
-
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings15 May 2024 0 repositories listed
-
Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer15 May 2024 0 repositories listed
-
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech10 May 2024 0 repositories listed
-
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition6 May 2024 0 repositories listed
-
Combining X-Vectors and Bayesian Batch Active Learning: Two-Stage Active Learning Pipeline for Speech Recognition3 May 2024 0 repositories listed
-
Efficient Compression of Multitask Multilingual Speech Models2 May 2024 0 repositories listed
-
Improving Membership Inference in ASR Model Auditing with Perturbed Loss Features2 May 2024 0 repositories listed
-
Sequence-to-sequence models in peer-to-peer learning: A practical application2 May 2024 0 repositories listed
-
Does Whisper understand Swiss German? An automatic, qualitative, and human evaluation30 Apr 2024 0 repositories listed
-
Automatic Speech Recognition System-Independent Word Error Rate Estimation25 Apr 2024 0 repositories listed
-
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF25 Apr 2024 0 repositories listed
-
Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm23 Apr 2024 0 repositories listed
-
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance23 Apr 2024 0 repositories listed
-
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech18 Apr 2024 0 repositories listed
-
Anatomy of Industrial Scale Multilingual ASR15 Apr 2024 0 repositories listed
-
12 Apr 2024 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task12 Apr 2024 0 repositories listed
-
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution11 Apr 2024 0 repositories listed
-
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping10 Apr 2024 0 repositories listed
-
The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge9 Apr 2024 0 repositories listed
-
Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian3 Apr 2024 0 repositories listed
-
Noise Masking Attacks and Defenses for Pretrained Speech Models2 Apr 2024 0 repositories listed
-
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models31 Mar 2024 0 repositories listed
-
LV-CTC: Non-autoregressive ASR with CTC and latent variable models28 Mar 2024 0 repositories listed
-
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition28 Mar 2024 0 repositories listed
-
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus27 Mar 2024 0 repositories listed
-
Extracting Biomedical Entities from Noisy Audio Transcripts26 Mar 2024 0 repositories listed
-
A Multimodal Approach to Device-Directed Speech Detection with Large Language Models21 Mar 2024 0 repositories listed
-
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech20 Mar 2024 0 repositories listed
-
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning20 Mar 2024 0 repositories listed
-
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition18 Mar 2024 0 repositories listed
-
Artificial Intelligence for Cochlear Implants: Review of Strategies, Challenges, and Perspectives17 Mar 2024 0 repositories listed
-
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children13 Mar 2024 0 repositories listed
-
The evaluation of a code-switched Sepedi-English automatic speech recognition system11 Mar 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.