Browse State-of-the-Art › Automatic Speech Recognition (ASR) › Papers, page 11
Automatic Speech Recognition (ASR)
Papers archive 2025-07-28
archive papers tagged: 3,012 · with a code link: 622 · where Syntology ran a sample: 77 (64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (77 of 3,012 tagged: 64 with a run with no instrument failure, 13 where every run was a failure of Syntology's instrument)
Page 11 of 31: papers 1,001 to 1,100 of 3,012, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Aligning Speech to Languages to Enhance Code-switching Speech Recognition9 Mar 2024 0 repositories listed
-
A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain7 Mar 2024 0 repositories listed
-
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition4 Mar 2024 0 repositories listed
-
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey2 Mar 2024 0 repositories listed
-
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview1 Mar 2024 0 repositories listed
-
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition29 Feb 2024 0 repositories listed
-
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems29 Feb 2024 0 repositories listed
-
Exploration of Adapter for Noise Robust Automatic Speech Recognition28 Feb 2024 0 repositories listed
-
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement27 Feb 2024 0 repositories listed
-
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models27 Feb 2024 0 repositories listed
-
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR21 Feb 2024 0 repositories listed
-
Ain't Misbehavin' -- Using LLMs to Generate Expressive Robot Behavior in Conversations with the Tabletop Robot Haru18 Feb 2024 0 repositories listed
-
The Balancing Act: Unmasking and Alleviating ASR Biases in Portuguese12 Feb 2024 0 repositories listed
-
The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models12 Feb 2024 0 repositories listed
-
Progressive unsupervised domain adaptation for ASR using ensemble models and multi-stage training7 Feb 2024 0 repositories listed
-
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems5 Feb 2024 0 repositories listed
-
Resolving Transcription Ambiguity in Spanish: A Hybrid Acoustic-Lexical System for Punctuation Restoration5 Feb 2024 0 repositories listed
-
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens3 Feb 2024 0 repositories listed
-
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents2 Feb 2024 0 repositories listed
-
Digits micro-model for accurate and secure transactions2 Feb 2024 0 repositories listed
-
Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges2 Feb 2024 0 repositories listed
-
Byte Pair Encoding Is All You Need For Automatic Bengali Speech Recognition28 Jan 2024 0 repositories listed
-
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline26 Jan 2024 0 repositories listed
-
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction24 Jan 2024 0 repositories listed
-
Locality enhanced dynamic biasing and sampling strategies for contextual ASR23 Jan 2024 0 repositories listed
-
Consistency Based Unsupervised Self-training For ASR Personalisation22 Jan 2024 0 repositories listed
-
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers22 Jan 2024 0 repositories listed
-
Using Large Language Model for End-to-End Chinese ASR and NER21 Jan 2024 0 repositories listed
-
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search19 Jan 2024 0 repositories listed
-
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition18 Jan 2024 0 repositories listed
-
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks18 Jan 2024 0 repositories listed
-
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition18 Jan 2024 0 repositories listed
-
Improving ASR Contextual Biasing with Guided Attention16 Jan 2024 0 repositories listed
-
Promptformer: Prompted Conformer Transducer for ASR14 Jan 2024 0 repositories listed
-
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization13 Jan 2024 0 repositories listed
-
Transcending Controlled Environments Assessing the Transferability of ASRRobust NLU Models to Real-World Applications12 Jan 2024 0 repositories listed
-
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese12 Jan 2024 0 repositories listed
-
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec211 Jan 2024 0 repositories listed
-
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction11 Jan 2024 0 repositories listed
-
Useful Blunders: Can Automated Speech Recognition Errors Improve Downstream Dementia Classification?10 Jan 2024 0 repositories listed
-
Continuously Learning New Words in Automatic Speech Recognition9 Jan 2024 0 repositories listed
-
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators8 Jan 2024 0 repositories listed
-
Exploratory Evaluation of Speech Content Masking8 Jan 2024 0 repositories listed
-
High-precision Voice Search Query Correction via Retrievable Speech-text Embedings8 Jan 2024 0 repositories listed
-
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR8 Jan 2024 0 repositories listed
-
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge7 Jan 2024 0 repositories listed
-
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition7 Jan 2024 0 repositories listed
-
Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models3 Jan 2024 0 repositories listed
-
Towards Probing Contact Center Large Language Models26 Dec 2023 0 repositories listed
-
Exploring data augmentation in bias mitigation against non-native-accented speech24 Dec 2023 0 repositories listed
-
BLSTM-Based Confidence Estimation for End-to-End Speech Recognition22 Dec 2023 0 repositories listed
-
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification22 Dec 2023 0 repositories listed
-
Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models20 Dec 2023 0 repositories listed
-
SpokesBiz -- an Open Corpus of Conversational Polish19 Dec 2023 0 repositories listed
-
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices16 Dec 2023 0 repositories listed
-
OAVA: the open audio-visual archives aggregator16 Dec 2023 0 repositories listed
-
LiteVSR: Efficient Visual Speech Recognition by Learning from Speech Representations of Unlabeled Data15 Dec 2023 0 repositories listed
-
FastInject: Injecting Unpaired Text Data into CTC-based ASR training14 Dec 2023 0 repositories listed
-
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models13 Dec 2023 0 repositories listed
-
Creating Spoken Dialog Systems in Ultra-Low Resourced Settings11 Dec 2023 0 repositories listed
-
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition6 Dec 2023 0 repositories listed
-
End-to-End Speech-to-Text Translation: A Survey2 Dec 2023 0 repositories listed
-
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data29 Nov 2023 0 repositories listed
-
How does end-to-end speech recognition training impact speech enhancement artifacts?20 Nov 2023 0 repositories listed
-
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition19 Nov 2023 0 repositories listed
-
ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding19 Nov 2023 0 repositories listed
-
Improving Large-scale Deep Biasing with Phoneme Features and Text-only Data in Streaming Transducer15 Nov 2023 0 repositories listed
-
Multi-channel Conversational Speaker Separation via Neural Diarization15 Nov 2023 0 repositories listed
-
Retrieve and Copy: Scaling ASR Personalization to Large Catalogs14 Nov 2023 0 repositories listed
-
On the Effectiveness of ASR Representations in Real-world Noisy Speech Emotion Recognition13 Nov 2023 0 repositories listed
-
1SPU: 1-step Speech Processing Unit8 Nov 2023 0 repositories listed
-
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning3 Nov 2023 0 repositories listed
-
Server-side Rescoring of Spoken Entity-centric Knowledge Queries for Virtual Assistants2 Nov 2023 0 repositories listed
-
RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios31 Oct 2023 0 repositories listed
-
Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation23 Oct 2023 0 repositories listed
-
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation22 Oct 2023 0 repositories listed
-
Intelligibility prediction with a pretrained noise-robust automatic speech recognition model20 Oct 2023 0 repositories listed
-
Unintended Memorization in Large ASR Models, and How to Mitigate It18 Oct 2023 0 repositories listed
-
Advanced accent/dialect identification and accentedness assessment with multi-embedding models and automatic speech recognition17 Oct 2023 0 repositories listed
-
Correction Focused Language Model Training for Speech Recognition17 Oct 2023 0 repositories listed
-
Generative error correction for code-switching speech recognition using large language models17 Oct 2023 0 repositories listed
-
Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition17 Oct 2023 0 repositories listed
-
VoxArabica: A Robust Dialect-Aware Arabic Speech Recognition System17 Oct 2023 0 repositories listed
-
Detecting Speech Abnormalities with a Perceiver-based Sequence Classifier that Leverages a Universal Speech Model16 Oct 2023 0 repositories listed
-
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring14 Oct 2023 0 repositories listed
-
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text12 Oct 2023 0 repositories listed
-
On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition12 Oct 2023 0 repositories listed
-
Acoustic Model Fusion for End-to-end Speech Recognition10 Oct 2023 0 repositories listed
-
Discriminative Speech Recognition Rescoring with Pre-trained Language Models10 Oct 2023 0 repositories listed
-
Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent Synthesis9 Oct 2023 0 repositories listed
-
Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech Recognition7 Oct 2023 0 repositories listed
-
A privacy-preserving method using secret key for convolutional neural network-based speech classification6 Oct 2023 0 repositories listed
-
An Integrated Algorithm for Robust and Imperceptible Audio Adversarial Examples5 Oct 2023 0 repositories listed
-
4 Oct 2023 0 repositories listed
-
One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition2 Oct 2023 0 repositories listed
-
AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR30 Sep 2023 0 repositories listed
-
AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition29 Sep 2023 0 repositories listed
-
Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm29 Sep 2023 0 repositories listed
-
Enhancing Code-switching Speech Recognition with Interactive Language Biases29 Sep 2023 0 repositories listed
-
29 Sep 2023 0 repositories listed Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.