Browse State-of-the-Art › Automatic Speech Recognition › Papers, page 14
Automatic Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 3,174 · with a code link: 677 · where Syntology ran a sample: 79 (62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (79 of 3,174 tagged: 62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 14 of 32: papers 1,301 to 1,400 of 3,174, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study27 Sep 2023 0 repositories listed
-
Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study27 Sep 2023 0 repositories listed
-
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference26 Sep 2023 0 repositories listed
-
Unsupervised Pre-Training for Vietnamese Automatic Speech Recognition in the HYKIST Project26 Sep 2023 0 repositories listed
-
AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data25 Sep 2023 0 repositories listed
-
Connecting Speech Encoder and Large Language Model for ASR25 Sep 2023 0 repositories listed
-
Cross-modal Alignment with Optimal Transport for CTC-based ASR24 Sep 2023 0 repositories listed
-
Speech enhancement with frequency domain auto-regressive modeling24 Sep 2023 0 repositories listed
-
My Science Tutor (MyST) -- A Large Corpus of Children's Conversational Speech23 Sep 2023 0 repositories listed
-
Affect Recognition in Conversations Using Large Language Models22 Sep 2023 0 repositories listed
-
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model22 Sep 2023 0 repositories listed
-
Importance of Smoothness Induced by Optimizers in FL4ASR: Towards Understanding Federated Learning for End-to-End ASR22 Sep 2023 0 repositories listed
-
Massive End-to-end Models for Short Search Queries22 Sep 2023 0 repositories listed
-
NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization22 Sep 2023 0 repositories listed
-
A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech Enhancement21 Sep 2023 0 repositories listed
-
Sparsely Shared LoRA on Whisper for Child Speech Recognition21 Sep 2023 0 repositories listed
-
AudioFool: Fast, Universal and synchronization-free Cross-Domain Attack on Speech Recognition20 Sep 2023 0 repositories listed
-
Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition20 Sep 2023 0 repositories listed
-
Exploring Speech Enhancement for Low-resource Speech Synthesis19 Sep 2023 0 repositories listed
-
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement19 Sep 2023 0 repositories listed
-
Semi-Autoregressive Streaming ASR With Label Context19 Sep 2023 0 repositories listed
-
A Multitask Training Approach to Enhance Whisper with Contextual Biasing and Open-Vocabulary Keyword Spotting18 Sep 2023 0 repositories listed
-
Corpus Synthesis for Zero-shot ASR domain Adaptation using Large Language Models18 Sep 2023 0 repositories listed
-
Distilling HuBERT with LSTMs via Decoupled Knowledge Distillation18 Sep 2023 0 repositories listed
-
HTEC: Human Transcription Error Correction18 Sep 2023 0 repositories listed
-
Instruction-Following Speech Recognition18 Sep 2023 0 repositories listed
-
Investigating End-to-End ASR Architectures for Long Form Audio Transcription18 Sep 2023 0 repositories listed
-
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning17 Sep 2023 0 repositories listed
-
Boosting End-to-End Multilingual Phoneme Recognition through Exploiting Universal Speech Attributes Constraints16 Sep 2023 0 repositories listed
-
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation16 Sep 2023 0 repositories listed
-
Improving Speech Recognition for African American English With Audio Classification16 Sep 2023 0 repositories listed
-
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription15 Sep 2023 0 repositories listed
-
t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability15 Sep 2023 0 repositories listed
-
Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network15 Sep 2023 0 repositories listed
-
CPPF: A contextual and post-processing-free model for automatic speech recognition14 Sep 2023 0 repositories listed
-
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks14 Sep 2023 0 repositories listed
-
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation14 Sep 2023 0 repositories listed
-
Can Whisper perform speech-based in-context learning?13 Sep 2023 0 repositories listed
-
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis13 Sep 2023 0 repositories listed
-
Open-vocabulary Keyword-spotting with Adaptive Instance Normalization13 Sep 2023 0 repositories listed
-
Improving Robustness of Neural Inverse Text Normalization via Data-Augmentation, Semi-Supervised Learning, and Post-Aligning Method12 Sep 2023 0 repositories listed
-
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults12 Sep 2023 0 repositories listed
-
Leveraging Large Language Models for Exploiting ASR Uncertainty9 Sep 2023 0 repositories listed
-
LanSER: Language-Model Supported Speech Emotion Recognition7 Sep 2023 0 repositories listed
-
Multiple Representation Transfer from Large Language Models to End-to-End ASR Systems7 Sep 2023 0 repositories listed
-
Bring the Noise: Introducing Noise Robustness to Pretrained Automatic Speech Recognition5 Sep 2023 0 repositories listed
-
TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-device ASR Models5 Sep 2023 0 repositories listed
-
AVATAR: Robust Voice Search Engine Leveraging Autoregressive Document Retrieval and Contrastive Learning4 Sep 2023 0 repositories listed
-
Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation4 Sep 2023 0 repositories listed
-
Contextual Biasing of Named-Entities with Large Language Models1 Sep 2023 0 repositories listed
-
Learning Speech Representation From Contrastive Token-Acoustic Pretraining1 Sep 2023 0 repositories listed
-
Knowledge Distillation from Non-streaming to Streaming ASR Encoder using Auxiliary Non-streaming Layer31 Aug 2023 0 repositories listed
-
ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers30 Aug 2023 0 repositories listed
-
Adapting Text-based Dialogue State Tracker for Spoken Dialogues29 Aug 2023 0 repositories listed
-
Neural approaches to spoken content embedding28 Aug 2023 0 repositories listed
-
Unsupervised Active Learning: Optimizing Labeling Cost-Effectiveness for Automatic Speech Recognition28 Aug 2023 0 repositories listed
-
Decoupled Structure for Improved Adaptability of End-to-End Models25 Aug 2023 0 repositories listed
-
Convoifilter: A case study of doing cocktail party speech recognition22 Aug 2023 0 repositories listed
-
Identifying depression-related topics in smartphone-collected free-response speech recordings using an automatic speech recognition system and a deep learning topic model22 Aug 2023 0 repositories listed
-
TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition21 Aug 2023 0 repositories listed
-
Indonesian Automatic Speech Recognition with XLSR-5320 Aug 2023 0 repositories listed
-
Accurate synthesis of Dysarthric Speech for ASR data augmentation16 Aug 2023 0 repositories listed
-
Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals16 Aug 2023 0 repositories listed
-
Improving CTC-AED model with integrated-CTC and auxiliary loss regularization15 Aug 2023 0 repositories listed
-
Text Injection for Capitalization and Turn-Taking Prediction in Speech Models14 Aug 2023 0 repositories listed
-
Using Text Injection to Improve Recognition of Personal Identifiers in Speech14 Aug 2023 0 repositories listed
-
Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition12 Aug 2023 0 repositories listed
-
Bilingual Streaming ASR with Grapheme units and Auxiliary Monolingual Loss11 Aug 2023 0 repositories listed
-
A Novel Self-training Approach for Low-resource Speech Recognition10 Aug 2023 0 repositories listed
-
Comparative Analysis of the wav2vec 2.0 Feature Extractor8 Aug 2023 0 repositories listed
-
Boosting Chinese ASR Error Correction with Dynamic Error Scaling Mechanism7 Aug 2023 0 repositories listed
-
ApproBiVT: Lead ASR Models to Generalize Better Using Approximated Bias-Variance Tradeoff Guided Early Stopping and Checkpoint Averaging5 Aug 2023 0 repositories listed
-
Federated Representation Learning for Automatic Speech Recognition3 Aug 2023 0 repositories listed
-
Careful Whisper -- leveraging advances in automatic speech recognition for robust and interpretable aphasia subtype classification2 Aug 2023 0 repositories listed
-
Pre-training End-to-end ASR Models with Augmented Speech Samples Queried by Text30 Jul 2023 0 repositories listed
-
Cascaded Cross-Modal Transformer for Request and Complaint Detection27 Jul 2023 0 repositories listed
-
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition26 Jul 2023 0 repositories listed
-
On-Device Speaker Anonymization of Acoustic Embeddings for ASR based onFlexible Location Gradient Reversal Layer25 Jul 2023 0 repositories listed
-
Boosting Punctuation Restoration with Data Generation and Reinforcement Learning24 Jul 2023 0 repositories listed
-
Integration of Frame- and Label-synchronous Beam Search for Streaming Encoder-decoder Speech Recognition24 Jul 2023 0 repositories listed
-
Robust Automatic Speech Recognition via WavAugment Guided Phoneme Adversarial Training24 Jul 2023 0 repositories listed
-
A meta learning scheme for fast accent domain expansion in Mandarin speech recognition23 Jul 2023 0 repositories listed
-
Exploring the Integration of Speech Separation and Recognition with Self-Supervised Learning Representation23 Jul 2023 0 repositories listed
-
Prompting Large Language Models with Speech Recognition Abilities21 Jul 2023 0 repositories listed
-
Model Adaptation for ASR in low-resource Indian Languages16 Jul 2023 0 repositories listed
-
Ed-Fed: A generic federated learning framework with resource-aware client selection for edge devices14 Jul 2023 0 repositories listed
-
Replay to Remember: Continual Layer-Specific Fine-tuning for German Speech Recognition14 Jul 2023 0 repositories listed
-
Representation Learning With Hidden Unit Clustering For Low Resource Speech Applications14 Jul 2023 0 repositories listed
-
Exploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical Study13 Jul 2023 0 repositories listed
-
Speech Diarization and ASR with GMM11 Jul 2023 0 repositories listed
-
Token-Level Serialized Output Training for Joint Streaming ASR and ST Leveraging Textual Alignments7 Jul 2023 0 repositories listed
-
Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture5 Jul 2023 0 repositories listed
-
Align With Purpose: Optimize Desired Properties in CTC Models with a General Plug-and-Play Framework4 Jul 2023 0 repositories listed
-
Boosting Norwegian Automatic Speech Recognition4 Jul 2023 0 repositories listed
-
Knowledge-Aware Audio-Grounded Generative Slot Filling for Limited Annotated Data4 Jul 2023 0 repositories listed
-
Transcribing Educational Videos Using Whisper: A preliminary study on using AI for transcribing educational videos4 Jul 2023 0 repositories listed
-
Multilingual Contextual Adapters To Improve Custom Word Recognition In Low-resource Languages3 Jul 2023 0 repositories listed
-
Conformer LLMs -- Convolution Augmented Large Language Models2 Jul 2023 0 repositories listed
-
Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters2 Jul 2023 0 repositories listed
-
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications29 Jun 2023 0 repositories listed