Browse State-of-the-Art › Automatic Speech Recognition › Papers, page 15
Automatic Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 3,174 · with a code link: 677 · where Syntology ran a sample: 79 (62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (79 of 3,174 tagged: 62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 15 of 32: papers 1,401 to 1,500 of 3,174, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Accelerating Transducers through Adjacent Token Merging28 Jun 2023 0 repositories listed
-
Master-ASR: Achieving Multilingual Scalability and Low-Resource Adaptation in ASR with Modular Learning23 Jun 2023 0 repositories listed
-
The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios23 Jun 2023 0 repositories listed
-
Exploring the Role of Audio in Video Captioning21 Jun 2023 0 repositories listed
-
Federated Self-Learning with Weak Supervision for Speech Recognition21 Jun 2023 0 repositories listed
-
Learning When to Trust Which Teacher for Weakly Supervised ASR21 Jun 2023 0 repositories listed
-
Mixture Encoder for Joint Speech Separation and Recognition21 Jun 2023 0 repositories listed
-
Strategies in Transfer Learning for Low-Resource Speech Synthesis: Phone Mapping, Features Input, and Source Language Selection21 Jun 2023 0 repositories listed
-
Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction15 Jun 2023 0 repositories listed
-
MobileASR: A resource-aware on-device learning framework for user voice personalization applications on mobile phones15 Jun 2023 0 repositories listed
-
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation14 Jun 2023 0 repositories listed
-
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR13 Jun 2023 0 repositories listed
-
Statistical Beamformer Exploiting Non-stationarity and Sparsity with Spatially Constrained ICA for Robust Speech Recognition13 Jun 2023 0 repositories listed
-
Multi-View Frequency-Attention Alternative to CNN Frontends for Automatic Speech Recognition12 Jun 2023 0 repositories listed
-
Multimodal Audio-textual Architecture for Robust Spoken Language Understanding12 Jun 2023 0 repositories listed
-
On the N-gram Approximation of Pre-trained Language Models12 Jun 2023 0 repositories listed
-
Impact of Experiencing Misrecognition by Teachable Agents on Learning and Rapport11 Jun 2023 0 repositories listed
-
What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model10 Jun 2023 0 repositories listed
-
Improving Frame-level Classifier for Word Timings with Non-peaky CTC in End-to-End Automatic Speech Recognition9 Jun 2023 0 repositories listed
-
Improving Language Model Integration for Neural Machine Translation8 Jun 2023 0 repositories listed
-
A study on the impact of Self-Supervised Learning on automatic dysarthric speech assessment7 Jun 2023 0 repositories listed
-
An ASR-Based Tutor for Learning to Read: How to Optimize Feedback to First Graders7 Jun 2023 0 repositories listed
-
FOOCTTS: Generating Arabic Speech with Acoustic Environment for Football Commentator7 Jun 2023 0 repositories listed
-
Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization7 Jun 2023 0 repositories listed
-
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses6 Jun 2023 0 repositories listed
-
Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics6 Jun 2023 0 repositories listed
-
Improving Fairness and Robustness in End-to-End Speech Recognition through unsupervised clustering6 Jun 2023 0 repositories listed
-
Incorporating L2 Phonemes Using Articulatory Features for Robust Speech Recognition5 Jun 2023 0 repositories listed
-
OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition5 Jun 2023 0 repositories listed
-
End-to-End Joint Target and Non-Target Speakers ASR4 Jun 2023 0 repositories listed
-
Audio-Visual Speech Enhancement with Score-Based Generative Models2 Jun 2023 0 repositories listed
-
Improved Training for End-to-End Streaming Automatic Speech Recognition Model with Punctuation2 Jun 2023 0 repositories listed
-
Streaming Speech-to-Confusion Network Speech Recognition2 Jun 2023 0 repositories listed
-
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication1 Jun 2023 0 repositories listed
-
AfriNames: Most ASR models "butcher" African Names1 Jun 2023 0 repositories listed
-
Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts1 Jun 2023 0 repositories listed
-
Encoder-decoder multimodal speaker change detection1 Jun 2023 0 repositories listed
-
Inspecting Spoken Language Understanding from Kids for Basic Math Learning at Home1 Jun 2023 0 repositories listed
-
Some voices are too common: Building fair speech recognition systems using the Common Voice dataset1 Jun 2023 0 repositories listed
-
Towards hate speech detection in low-resource languages: Comparing ASR to acoustic word embeddings on Wolof and Swahili1 Jun 2023 0 repositories listed
-
Accurate and Structured Pruning for Efficient Automatic Speech Recognition31 May 2023 0 repositories listed
-
Simple yet Effective Code-Switching Language Identification with Multitask Pre-Training and Transfer Learning31 May 2023 0 repositories listed
-
Strategies for improving low resource speech to text translation relying on pre-trained ASR models31 May 2023 0 repositories listed
-
VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition31 May 2023 0 repositories listed
-
Zero-Shot Automatic Pronunciation Assessment31 May 2023 0 repositories listed
-
Adapting Multi-Lingual ASR Models for Handling Multiple Talkers30 May 2023 0 repositories listed
-
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions30 May 2023 0 repositories listed
-
Towards Selection of Text-to-speech Data to Augment ASR Training30 May 2023 0 repositories listed
-
29 May 2023 0 repositories listed
-
Building Accurate Low Latency ASR for Streaming Voice Search29 May 2023 0 repositories listed
-
Can We Trust Explainable AI Methods on ASR? An Evaluation on Phoneme Recognition29 May 2023 0 repositories listed
-
Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target29 May 2023 0 repositories listed
-
Retraining-free Customized ASR for Enharmonic Words Based on a Named-Entity-Aware Model and Phoneme Similarity Estimation29 May 2023 0 repositories listed
-
2-bit Conformer quantization for automatic speech recognition26 May 2023 0 repositories listed
-
DisfluencyFixer: A tool to enhance Language Learning through Speech To Speech Disfluency Correction26 May 2023 0 repositories listed
-
ASR and Emotional Speech: A Word-Level Investigation of the Mutual Impact of Speech and Emotion Recognition25 May 2023 0 repositories listed
-
Improving Scheduled Sampling for Neural Transducer-based ASR25 May 2023 0 repositories listed
-
INTapt: Information-Theoretic Adversarial Prompt Tuning for Enhanced Non-Native Speech Recognition25 May 2023 0 repositories listed
-
Mixture-of-Expert Conformer for Streaming Multilingual ASR25 May 2023 0 repositories listed
-
Svarah: Evaluating English ASR Systems on Indian Accents25 May 2023 0 repositories listed
-
Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator25 May 2023 0 repositories listed
-
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation24 May 2023 0 repositories listed
-
InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition24 May 2023 0 repositories listed
-
Iteratively Improving Speech Recognition and Voice Conversion24 May 2023 0 repositories listed
-
BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR23 May 2023 0 repositories listed
-
Cross-lingual Knowledge Transfer and Iterative Pseudo-labeling for Low-Resource Speech Recognition with Transducers23 May 2023 0 repositories listed
-
Evaluating OpenAI's Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person23 May 2023 0 repositories listed
-
Graph Meets LLM: A Novel Approach to Collaborative Filtering for Robust Conversational Understanding23 May 2023 0 repositories listed
-
On the Transferability of Whisper-based Representations for "In-the-Wild" Cross-Task Downstream Speech Applications23 May 2023 0 repositories listed
-
Personalized Predictive ASR for Latency Reduction in Voice Assistants23 May 2023 0 repositories listed
-
SE-Bridge: Speech Enhancement with Consistent Brownian Bridge23 May 2023 0 repositories listed
-
TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition23 May 2023 0 repositories listed
-
Debiased Automatic Speech Recognition for Dysarthric Speech via Sample Reweighting with Sample Affinity Test22 May 2023 0 repositories listed
-
GNCformer Enhanced Self-attention for Automatic Speech Recognition22 May 2023 0 repositories listed
-
Text Generation with Speech Synthesis for ASR Data Augmentation22 May 2023 0 repositories listed
-
CASA-ASR: Context-Aware Speaker-Attributed ASR21 May 2023 0 repositories listed
-
Hystoc: Obtaining word confidences for fusion of end-to-end ASR systems21 May 2023 0 repositories listed
-
On the Efficacy and Noise-Robustness of Jointly Learned Speech Emotion and Automatic Speech Recognition21 May 2023 0 repositories listed
-
Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction21 May 2023 0 repositories listed
-
VAKTA-SETU: A Speech-to-Speech Machine Translation Service in Select Indic Languages21 May 2023 0 repositories listed
-
Self-supervised representations in speech-based depression detection20 May 2023 0 repositories listed
-
Unsupervised ASR via Cross-Lingual Pseudo-Labeling19 May 2023 0 repositories listed
-
A Lexical-aware Non-autoregressive Transformer-based ASR Model18 May 2023 0 repositories listed
-
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark18 May 2023 0 repositories listed
-
Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation Based Voice Conversion16 May 2023 0 repositories listed
-
Application-Agnostic Language Modeling for On-Device ASR16 May 2023 0 repositories listed
-
Critical Appraisal of Artificial Intelligence-Mediated Communication15 May 2023 0 repositories listed
-
OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking15 May 2023 0 repositories listed
-
Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations14 May 2023 0 repositories listed
-
Continual Learning for End-to-End ASR by Averaging Domain Experts12 May 2023 0 repositories listed
-
Investigating the Sensitivity of Automatic Speech Recognition Systems to Phonetic Variation in L2 Englishes12 May 2023 0 repositories listed
-
Masked Audio Text Encoders are Effective Multi-Modal Rescorers11 May 2023 0 repositories listed
-
Quran Recitation Recognition using End-to-End Deep Learning10 May 2023 0 repositories listed
-
Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models9 May 2023 0 repositories listed
-
Robust Acoustic and Semantic Contextual Biasing in Neural Transducers for Speech Recognition9 May 2023 0 repositories listed
-
Who Needs Decoders? Efficient Estimation of Sequence-level Attributes9 May 2023 0 repositories listed
-
8 May 2023 0 repositories listed
-
Multi-Temporal Lip-Audio Memory for Visual Speech Recognition8 May 2023 0 repositories listed
-
Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers7 May 2023 0 repositories listed
-
Employing Hybrid Deep Neural Networks on Dari Speech4 May 2023 0 repositories listed