Browse State-of-the-Art › Automatic Speech Recognition (ASR) › Papers, page 12
Automatic Speech Recognition (ASR)
Papers archive 2025-07-28
archive papers tagged: 3,012 · with a code link: 622 · where Syntology ran a sample: 77 (64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (77 of 3,012 tagged: 64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument)
Page 12 of 31: papers 1,101 to 1,200 of 3,012, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition29 Sep 2023 0 repositories listed
-
The Gift of Feedback: Improving ASR Model Quality by Learning from User Corrections through Federated Learning29 Sep 2023 0 repositories listed
-
Wiki-En-ASR-Adapt: Large-scale synthetic dataset for English ASR Customization29 Sep 2023 0 repositories listed
-
Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR28 Sep 2023 0 repositories listed
-
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference26 Sep 2023 0 repositories listed
-
Connecting Speech Encoder and Large Language Model for ASR25 Sep 2023 0 repositories listed
-
Cross-modal Alignment with Optimal Transport for CTC-based ASR24 Sep 2023 0 repositories listed
-
Speech enhancement with frequency domain auto-regressive modeling24 Sep 2023 0 repositories listed
-
Affect Recognition in Conversations Using Large Language Models22 Sep 2023 0 repositories listed
-
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model22 Sep 2023 0 repositories listed
-
Importance of Smoothness Induced by Optimizers in FL4ASR: Towards Understanding Federated Learning for End-to-End ASR22 Sep 2023 0 repositories listed
-
Massive End-to-end Models for Short Search Queries22 Sep 2023 0 repositories listed
-
Sparsely Shared LoRA on Whisper for Child Speech Recognition21 Sep 2023 0 repositories listed
-
Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition20 Sep 2023 0 repositories listed
-
Exploring Speech Enhancement for Low-resource Speech Synthesis19 Sep 2023 0 repositories listed
-
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement19 Sep 2023 0 repositories listed
-
Semi-Autoregressive Streaming ASR With Label Context19 Sep 2023 0 repositories listed
-
A Multitask Training Approach to Enhance Whisper with Contextual Biasing and Open-Vocabulary Keyword Spotting18 Sep 2023 0 repositories listed
-
Corpus Synthesis for Zero-shot ASR domain Adaptation using Large Language Models18 Sep 2023 0 repositories listed
-
HTEC: Human Transcription Error Correction18 Sep 2023 0 repositories listed
-
Instruction-Following Speech Recognition18 Sep 2023 0 repositories listed
-
Investigating End-to-End ASR Architectures for Long Form Audio Transcription18 Sep 2023 0 repositories listed
-
Boosting End-to-End Multilingual Phoneme Recognition through Exploiting Universal Speech Attributes Constraints16 Sep 2023 0 repositories listed
-
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation16 Sep 2023 0 repositories listed
-
Improving Speech Recognition for African American English With Audio Classification16 Sep 2023 0 repositories listed
-
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription15 Sep 2023 0 repositories listed
-
t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability15 Sep 2023 0 repositories listed
-
Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network15 Sep 2023 0 repositories listed
-
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks14 Sep 2023 0 repositories listed
-
Can Whisper perform speech-based in-context learning?13 Sep 2023 0 repositories listed
-
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis13 Sep 2023 0 repositories listed
-
Open-vocabulary Keyword-spotting with Adaptive Instance Normalization13 Sep 2023 0 repositories listed
-
Improving Robustness of Neural Inverse Text Normalization via Data-Augmentation, Semi-Supervised Learning, and Post-Aligning Method12 Sep 2023 0 repositories listed
-
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults12 Sep 2023 0 repositories listed
-
Leveraging Large Language Models for Exploiting ASR Uncertainty9 Sep 2023 0 repositories listed
-
Multiple Representation Transfer from Large Language Models to End-to-End ASR Systems7 Sep 2023 0 repositories listed
-
Bring the Noise: Introducing Noise Robustness to Pretrained Automatic Speech Recognition5 Sep 2023 0 repositories listed
-
TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-device ASR Models5 Sep 2023 0 repositories listed
-
AVATAR: Robust Voice Search Engine Leveraging Autoregressive Document Retrieval and Contrastive Learning4 Sep 2023 0 repositories listed
-
Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation4 Sep 2023 0 repositories listed
-
Contextual Biasing of Named-Entities with Large Language Models1 Sep 2023 0 repositories listed
-
Learning Speech Representation From Contrastive Token-Acoustic Pretraining1 Sep 2023 0 repositories listed
-
Knowledge Distillation from Non-streaming to Streaming ASR Encoder using Auxiliary Non-streaming Layer31 Aug 2023 0 repositories listed
-
ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers30 Aug 2023 0 repositories listed
-
Unsupervised Active Learning: Optimizing Labeling Cost-Effectiveness for Automatic Speech Recognition28 Aug 2023 0 repositories listed
-
Decoupled Structure for Improved Adaptability of End-to-End Models25 Aug 2023 0 repositories listed
-
Convoifilter: A case study of doing cocktail party speech recognition22 Aug 2023 0 repositories listed
-
TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition21 Aug 2023 0 repositories listed
-
Indonesian Automatic Speech Recognition with XLSR-5320 Aug 2023 0 repositories listed
-
Accurate synthesis of Dysarthric Speech for ASR data augmentation16 Aug 2023 0 repositories listed
-
Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals16 Aug 2023 0 repositories listed
-
Improving CTC-AED model with integrated-CTC and auxiliary loss regularization15 Aug 2023 0 repositories listed
-
Text Injection for Capitalization and Turn-Taking Prediction in Speech Models14 Aug 2023 0 repositories listed
-
Using Text Injection to Improve Recognition of Personal Identifiers in Speech14 Aug 2023 0 repositories listed
-
Bilingual Streaming ASR with Grapheme units and Auxiliary Monolingual Loss11 Aug 2023 0 repositories listed
-
A Novel Self-training Approach for Low-resource Speech Recognition10 Aug 2023 0 repositories listed
-
Comparative Analysis of the wav2vec 2.0 Feature Extractor8 Aug 2023 0 repositories listed
-
Boosting Chinese ASR Error Correction with Dynamic Error Scaling Mechanism7 Aug 2023 0 repositories listed
-
ApproBiVT: Lead ASR Models to Generalize Better Using Approximated Bias-Variance Tradeoff Guided Early Stopping and Checkpoint Averaging5 Aug 2023 0 repositories listed
-
Cascaded Cross-Modal Transformer for Request and Complaint Detection27 Jul 2023 0 repositories listed
-
On-Device Speaker Anonymization of Acoustic Embeddings for ASR based onFlexible Location Gradient Reversal Layer25 Jul 2023 0 repositories listed
-
Boosting Punctuation Restoration with Data Generation and Reinforcement Learning24 Jul 2023 0 repositories listed
-
Robust Automatic Speech Recognition via WavAugment Guided Phoneme Adversarial Training24 Jul 2023 0 repositories listed
-
A meta learning scheme for fast accent domain expansion in Mandarin speech recognition23 Jul 2023 0 repositories listed
-
Exploring the Integration of Speech Separation and Recognition with Self-Supervised Learning Representation23 Jul 2023 0 repositories listed
-
Prompting Large Language Models with Speech Recognition Abilities21 Jul 2023 0 repositories listed
-
Model Adaptation for ASR in low-resource Indian Languages16 Jul 2023 0 repositories listed
-
Replay to Remember: Continual Layer-Specific Fine-tuning for German Speech Recognition14 Jul 2023 0 repositories listed
-
Representation Learning With Hidden Unit Clustering For Low Resource Speech Applications14 Jul 2023 0 repositories listed
-
Exploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical Study13 Jul 2023 0 repositories listed
-
Speech Diarization and ASR with GMM11 Jul 2023 0 repositories listed
-
Token-Level Serialized Output Training for Joint Streaming ASR and ST Leveraging Textual Alignments7 Jul 2023 0 repositories listed
-
Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture5 Jul 2023 0 repositories listed
-
Align With Purpose: Optimize Desired Properties in CTC Models with a General Plug-and-Play Framework4 Jul 2023 0 repositories listed
-
Boosting Norwegian Automatic Speech Recognition4 Jul 2023 0 repositories listed
-
Knowledge-Aware Audio-Grounded Generative Slot Filling for Limited Annotated Data4 Jul 2023 0 repositories listed
-
Transcribing Educational Videos Using Whisper: A preliminary study on using AI for transcribing educational videos4 Jul 2023 0 repositories listed
-
Multilingual Contextual Adapters To Improve Custom Word Recognition In Low-resource Languages3 Jul 2023 0 repositories listed
-
Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters2 Jul 2023 0 repositories listed
-
Accelerating Transducers through Adjacent Token Merging28 Jun 2023 0 repositories listed
-
Master-ASR: Achieving Multilingual Scalability and Low-Resource Adaptation in ASR with Modular Learning23 Jun 2023 0 repositories listed
-
The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios23 Jun 2023 0 repositories listed
-
Exploring the Role of Audio in Video Captioning21 Jun 2023 0 repositories listed
-
Federated Self-Learning with Weak Supervision for Speech Recognition21 Jun 2023 0 repositories listed
-
Learning When to Trust Which Teacher for Weakly Supervised ASR21 Jun 2023 0 repositories listed
-
Mixture Encoder for Joint Speech Separation and Recognition21 Jun 2023 0 repositories listed
-
Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction15 Jun 2023 0 repositories listed
-
MobileASR: A resource-aware on-device learning framework for user voice personalization applications on mobile phones15 Jun 2023 0 repositories listed
-
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation14 Jun 2023 0 repositories listed
-
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR13 Jun 2023 0 repositories listed
-
Statistical Beamformer Exploiting Non-stationarity and Sparsity with Spatially Constrained ICA for Robust Speech Recognition13 Jun 2023 0 repositories listed
-
Multimodal Audio-textual Architecture for Robust Spoken Language Understanding12 Jun 2023 0 repositories listed
-
On the N-gram Approximation of Pre-trained Language Models12 Jun 2023 0 repositories listed
-
Impact of Experiencing Misrecognition by Teachable Agents on Learning and Rapport11 Jun 2023 0 repositories listed
-
Improving Frame-level Classifier for Word Timings with Non-peaky CTC in End-to-End Automatic Speech Recognition9 Jun 2023 0 repositories listed
-
A study on the impact of Self-Supervised Learning on automatic dysarthric speech assessment7 Jun 2023 0 repositories listed
-
An ASR-Based Tutor for Learning to Read: How to Optimize Feedback to First Graders7 Jun 2023 0 repositories listed
-
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses6 Jun 2023 0 repositories listed
-
Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics6 Jun 2023 0 repositories listed
-
Improving Fairness and Robustness in End-to-End Speech Recognition through unsupervised clustering6 Jun 2023 0 repositories listed