Browse State-of-the-Art › Automatic Speech Recognition (ASR) › Papers, page 13
Automatic Speech Recognition (ASR)
Papers archive 2025-07-28
archive papers tagged: 3,012 · with a code link: 622 · where Syntology ran a sample: 77 (64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (77 of 3,012 tagged: 64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument)
Page 13 of 31: papers 1,201 to 1,300 of 3,012, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Incorporating L2 Phonemes Using Articulatory Features for Robust Speech Recognition5 Jun 2023 0 repositories listed
-
OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition5 Jun 2023 0 repositories listed
-
End-to-End Joint Target and Non-Target Speakers ASR4 Jun 2023 0 repositories listed
-
Streaming Speech-to-Confusion Network Speech Recognition2 Jun 2023 0 repositories listed
-
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication1 Jun 2023 0 repositories listed
-
AfriNames: Most ASR models "butcher" African Names1 Jun 2023 0 repositories listed
-
Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts1 Jun 2023 0 repositories listed
-
Inspecting Spoken Language Understanding from Kids for Basic Math Learning at Home1 Jun 2023 0 repositories listed
-
Some voices are too common: Building fair speech recognition systems using the Common Voice dataset1 Jun 2023 0 repositories listed
-
Towards hate speech detection in low-resource languages: Comparing ASR to acoustic word embeddings on Wolof and Swahili1 Jun 2023 0 repositories listed
-
Accurate and Structured Pruning for Efficient Automatic Speech Recognition31 May 2023 0 repositories listed
-
Simple yet Effective Code-Switching Language Identification with Multitask Pre-Training and Transfer Learning31 May 2023 0 repositories listed
-
VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition31 May 2023 0 repositories listed
-
Zero-Shot Automatic Pronunciation Assessment31 May 2023 0 repositories listed
-
Adapting Multi-Lingual ASR Models for Handling Multiple Talkers30 May 2023 0 repositories listed
-
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions30 May 2023 0 repositories listed
-
Towards Selection of Text-to-speech Data to Augment ASR Training30 May 2023 0 repositories listed
-
Building Accurate Low Latency ASR for Streaming Voice Search29 May 2023 0 repositories listed
-
Can We Trust Explainable AI Methods on ASR? An Evaluation on Phoneme Recognition29 May 2023 0 repositories listed
-
Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target29 May 2023 0 repositories listed
-
2-bit Conformer quantization for automatic speech recognition26 May 2023 0 repositories listed
-
DisfluencyFixer: A tool to enhance Language Learning through Speech To Speech Disfluency Correction26 May 2023 0 repositories listed
-
ASR and Emotional Speech: A Word-Level Investigation of the Mutual Impact of Speech and Emotion Recognition25 May 2023 0 repositories listed
-
Improving Scheduled Sampling for Neural Transducer-based ASR25 May 2023 0 repositories listed
-
INTapt: Information-Theoretic Adversarial Prompt Tuning for Enhanced Non-Native Speech Recognition25 May 2023 0 repositories listed
-
Svarah: Evaluating English ASR Systems on Indian Accents25 May 2023 0 repositories listed
-
Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator25 May 2023 0 repositories listed
-
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation24 May 2023 0 repositories listed
-
InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition24 May 2023 0 repositories listed
-
Iteratively Improving Speech Recognition and Voice Conversion24 May 2023 0 repositories listed
-
BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR23 May 2023 0 repositories listed
-
Evaluating OpenAI's Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person23 May 2023 0 repositories listed
-
Graph Meets LLM: A Novel Approach to Collaborative Filtering for Robust Conversational Understanding23 May 2023 0 repositories listed
-
On the Transferability of Whisper-based Representations for "In-the-Wild" Cross-Task Downstream Speech Applications23 May 2023 0 repositories listed
-
Personalized Predictive ASR for Latency Reduction in Voice Assistants23 May 2023 0 repositories listed
-
SE-Bridge: Speech Enhancement with Consistent Brownian Bridge23 May 2023 0 repositories listed
-
TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition23 May 2023 0 repositories listed
-
GNCformer Enhanced Self-attention for Automatic Speech Recognition22 May 2023 0 repositories listed
-
Text Generation with Speech Synthesis for ASR Data Augmentation22 May 2023 0 repositories listed
-
On the Efficacy and Noise-Robustness of Jointly Learned Speech Emotion and Automatic Speech Recognition21 May 2023 0 repositories listed
-
Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction21 May 2023 0 repositories listed
-
VAKTA-SETU: A Speech-to-Speech Machine Translation Service in Select Indic Languages21 May 2023 0 repositories listed
-
Self-supervised representations in speech-based depression detection20 May 2023 0 repositories listed
-
Unsupervised ASR via Cross-Lingual Pseudo-Labeling19 May 2023 0 repositories listed
-
A Lexical-aware Non-autoregressive Transformer-based ASR Model18 May 2023 0 repositories listed
-
Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation Based Voice Conversion16 May 2023 0 repositories listed
-
Critical Appraisal of Artificial Intelligence-Mediated Communication15 May 2023 0 repositories listed
-
OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking15 May 2023 0 repositories listed
-
Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations14 May 2023 0 repositories listed
-
Investigating the Sensitivity of Automatic Speech Recognition Systems to Phonetic Variation in L2 Englishes12 May 2023 0 repositories listed
-
Masked Audio Text Encoders are Effective Multi-Modal Rescorers11 May 2023 0 repositories listed
-
Quran Recitation Recognition using End-to-End Deep Learning10 May 2023 0 repositories listed
-
Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models9 May 2023 0 repositories listed
-
Who Needs Decoders? Efficient Estimation of Sequence-level Attributes9 May 2023 0 repositories listed
-
Multi-Temporal Lip-Audio Memory for Visual Speech Recognition8 May 2023 0 repositories listed
-
Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers7 May 2023 0 repositories listed
-
Employing Hybrid Deep Neural Networks on Dari Speech4 May 2023 0 repositories listed
-
End-to-end spoken language understanding using joint CTC loss and self-supervised, pretrained acoustic encoders4 May 2023 0 repositories listed
-
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks4 May 2023 0 repositories listed
-
A Study on the Integration of Pipeline and E2E SLU systems for Spoken Semantic Parsing toward STOP Quality Challenge2 May 2023 0 repositories listed
-
Building a Non-native Speech Corpus Featuring Chinese-English Bilingual Children: Compilation and Rationale30 Apr 2023 0 repositories listed
-
Deep Transfer Learning for Automatic Speech Recognition: Towards Better Generalization27 Apr 2023 0 repositories listed
-
Understanding Shared Speech-Text Representations27 Apr 2023 0 repositories listed
-
Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition24 Apr 2023 0 repositories listed
-
Non-autoregressive End-to-end Approaches for Joint Automatic Speech Recognition and Spoken Language Understanding21 Apr 2023 0 repositories listed
-
Towards the Universal Defense for Query-Based Audio Adversarial Attacks20 Apr 2023 0 repositories listed
-
Security and Privacy Problems in Voice Assistant Applications: A Survey19 Apr 2023 0 repositories listed
-
Multimodal Short Video Rumor Detection System Based on Contrastive Learning17 Apr 2023 0 repositories listed
-
A Virtual Simulation-Pilot Agent for Training of Air Traffic Controllers16 Apr 2023 0 repositories listed
-
A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition15 Apr 2023 0 repositories listed
-
Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training12 Apr 2023 0 repositories listed
-
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR11 Apr 2023 0 repositories listed
-
Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data4 Apr 2023 0 repositories listed
-
Self-Supervised Learning-Based Source Separation for Meeting Data3 Apr 2023 0 repositories listed
-
31 Mar 2023 0 repositories listed
-
31 Mar 2023 0 repositories listed
-
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers30 Mar 2023 0 repositories listed
-
Joint unsupervised and supervised learning for context-aware language identification29 Mar 2023 0 repositories listed
-
Text is All You Need: Personalizing ASR Models using Controllable Speech Synthesis27 Mar 2023 0 repositories listed
-
23 Mar 2023 0 repositories listed
-
Enhancing Unsupervised Speech Recognition with Diffusion GANs23 Mar 2023 0 repositories listed
-
Code-Switching Text Generation and Injection in Mandarin-English ASR20 Mar 2023 0 repositories listed
-
Knowledge Distillation from Multiple Foundation Models for End-to-End Speech Recognition20 Mar 2023 0 repositories listed
-
A Deep Learning System for Domain-specific Speech Recognition18 Mar 2023 0 repositories listed
-
DistillW2V2: A Small and Streaming Wav2vec 2.0 Based ASR Model16 Mar 2023 0 repositories listed
-
Visual Information Matters for ASR Error Correction16 Mar 2023 0 repositories listed
-
Improving Accented Speech Recognition with Multi-Domain Training14 Mar 2023 0 repositories listed
-
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings10 Mar 2023 0 repositories listed
-
MIXPGD: Hybrid Adversarial Training for Speech Recognition Systems10 Mar 2023 0 repositories listed
-
wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts6 Mar 2023 0 repositories listed
-
End-to-End Speech Recognition: A Survey3 Mar 2023 0 repositories listed
-
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages2 Mar 2023 0 repositories listed
-
Leveraging Large Text Corpora for End-to-End Speech Summarization2 Mar 2023 0 repositories listed
-
Leveraging Redundancy in Multiple Audio Signals for Far-Field Speech Recognition1 Mar 2023 0 repositories listed
-
N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space1 Mar 2023 0 repositories listed
-
A Comparison of Speech Data Augmentation Methods Using S3PRL Toolkit27 Feb 2023 0 repositories listed
-
Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video27 Feb 2023 0 repositories listed
-
Diacritic Recognition Performance in Arabic ASR27 Feb 2023 0 repositories listed
-
Improving Medical Speech-to-Text Accuracy with Vision-Language Pre-training Model27 Feb 2023 0 repositories listed
-
MoLE : Mixture of Language Experts for Multi-Lingual Automatic Speech Recognition27 Feb 2023 0 repositories listed