Browse State-of-the-Art › Automatic Speech Recognition › Papers, page 13
Automatic Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 3,174 · with a code link: 677 · where Syntology ran a sample: 79 (62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (79 of 3,174 tagged: 62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 13 of 32: papers 1,201 to 1,300 of 3,174, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
On Speech Pre-emphasis as a Simple and Inexpensive Method to Boost Speech Enhancement17 Jan 2024 0 repositories listed
-
Improving ASR Contextual Biasing with Guided Attention16 Jan 2024 0 repositories listed
-
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization16 Jan 2024 0 repositories listed
-
NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription16 Jan 2024 0 repositories listed
-
SeMaScore : a new evaluation metric for automatic speech recognition tasks15 Jan 2024 0 repositories listed
-
Promptformer: Prompted Conformer Transducer for ASR14 Jan 2024 0 repositories listed
-
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization13 Jan 2024 0 repositories listed
-
Transcending Controlled Environments Assessing the Transferability of ASRRobust NLU Models to Real-World Applications12 Jan 2024 0 repositories listed
-
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese12 Jan 2024 0 repositories listed
-
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec211 Jan 2024 0 repositories listed
-
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction11 Jan 2024 0 repositories listed
-
Useful Blunders: Can Automated Speech Recognition Errors Improve Downstream Dementia Classification?10 Jan 2024 0 repositories listed
-
Continuously Learning New Words in Automatic Speech Recognition9 Jan 2024 0 repositories listed
-
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators8 Jan 2024 0 repositories listed
-
Exploratory Evaluation of Speech Content Masking8 Jan 2024 0 repositories listed
-
High-precision Voice Search Query Correction via Retrievable Speech-text Embedings8 Jan 2024 0 repositories listed
-
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR8 Jan 2024 0 repositories listed
-
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge7 Jan 2024 0 repositories listed
-
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition7 Jan 2024 0 repositories listed
-
Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models3 Jan 2024 0 repositories listed
-
The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge26 Dec 2023 0 repositories listed
-
Towards Probing Contact Center Large Language Models26 Dec 2023 0 repositories listed
-
Exploring data augmentation in bias mitigation against non-native-accented speech24 Dec 2023 0 repositories listed
-
BLSTM-Based Confidence Estimation for End-to-End Speech Recognition22 Dec 2023 0 repositories listed
-
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification22 Dec 2023 0 repositories listed
-
Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models20 Dec 2023 0 repositories listed
-
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?19 Dec 2023 0 repositories listed
-
SpokesBiz -- an Open Corpus of Conversational Polish19 Dec 2023 0 repositories listed
-
Efficiency-oriented approaches for self-supervised speech representation learning18 Dec 2023 0 repositories listed
-
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices16 Dec 2023 0 repositories listed
-
OAVA: the open audio-visual archives aggregator16 Dec 2023 0 repositories listed
-
Generative Context-aware Fine-tuning of Self-supervised Speech Models15 Dec 2023 0 repositories listed
-
Leveraging Language ID to Calculate Intermediate CTC Loss for Enhanced Code-Switching Speech Recognition15 Dec 2023 0 repositories listed
-
LiteVSR: Efficient Visual Speech Recognition by Learning from Speech Representations of Unlabeled Data15 Dec 2023 0 repositories listed
-
Audio-visual fine-tuning of audio-only ASR models14 Dec 2023 0 repositories listed
-
FastInject: Injecting Unpaired Text Data into CTC-based ASR training14 Dec 2023 0 repositories listed
-
PhasePerturbation: Speech Data Augmentation via Phase Perturbation for Automatic Speech Recognition13 Dec 2023 0 repositories listed
-
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models13 Dec 2023 0 repositories listed
-
Self-supervised Adaptive Pre-training of Multilingual Speech Models for Language and Dialect Identification12 Dec 2023 0 repositories listed
-
Creating Spoken Dialog Systems in Ultra-Low Resourced Settings11 Dec 2023 0 repositories listed
-
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition6 Dec 2023 0 repositories listed
-
Multimodal Data and Resource Efficient Device-Directed Speech Detection with Large Foundation Models6 Dec 2023 0 repositories listed
-
End-to-End Speech-to-Text Translation: A Survey2 Dec 2023 0 repositories listed
-
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data29 Nov 2023 0 repositories listed
-
Weak Alignment Supervision from Hybrid Model Improves End-to-end ASR24 Nov 2023 0 repositories listed
-
Soft Random Sampling: A Theoretical and Empirical Analysis21 Nov 2023 0 repositories listed
-
App for Resume-Based Job Matching with Speech Interviews and Grammar Analysis: A Review20 Nov 2023 0 repositories listed
-
How does end-to-end speech recognition training impact speech enhancement artifacts?20 Nov 2023 0 repositories listed
-
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition19 Nov 2023 0 repositories listed
-
ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding19 Nov 2023 0 repositories listed
-
Improving Large-scale Deep Biasing with Phoneme Features and Text-only Data in Streaming Transducer15 Nov 2023 0 repositories listed
-
Multi-channel Conversational Speaker Separation via Neural Diarization15 Nov 2023 0 repositories listed
-
Retrieve and Copy: Scaling ASR Personalization to Large Catalogs14 Nov 2023 0 repositories listed
-
On the Effectiveness of ASR Representations in Real-world Noisy Speech Emotion Recognition13 Nov 2023 0 repositories listed
-
1SPU: 1-step Speech Processing Unit8 Nov 2023 0 repositories listed
-
Fine-tuning convergence model in Bengali speech recognition7 Nov 2023 0 repositories listed
-
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning3 Nov 2023 0 repositories listed
-
Server-side Rescoring of Spoken Entity-centric Knowledge Queries for Virtual Assistants2 Nov 2023 0 repositories listed
-
RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios31 Oct 2023 0 repositories listed
-
Combining Language Models For Specialized Domains: A Colorful Approach30 Oct 2023 0 repositories listed
-
Dialect Adaptation and Data Augmentation for Low-Resource ASR: TalTech Systems for the MADASR 2023 Challenge26 Oct 2023 0 repositories listed
-
Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation23 Oct 2023 0 repositories listed
-
Modality Dropout for Multimodal Device Directed Speech Detection using Verbal and Non-Verbal Features23 Oct 2023 0 repositories listed
-
Quantifying the Dialect Gap and its Correlates Across Languages23 Oct 2023 0 repositories listed
-
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation22 Oct 2023 0 repositories listed
-
Intelligibility prediction with a pretrained noise-robust automatic speech recognition model20 Oct 2023 0 repositories listed
-
The CHiME-7 Challenge: System Description and Performance of NeMo Team's DASR System18 Oct 2023 0 repositories listed
-
Unintended Memorization in Large ASR Models, and How to Mitigate It18 Oct 2023 0 repositories listed
-
Advanced accent/dialect identification and accentedness assessment with multi-embedding models and automatic speech recognition17 Oct 2023 0 repositories listed
-
Correction Focused Language Model Training for Speech Recognition17 Oct 2023 0 repositories listed
-
Generative error correction for code-switching speech recognition using large language models17 Oct 2023 0 repositories listed
-
Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition17 Oct 2023 0 repositories listed
-
VoxArabica: A Robust Dialect-Aware Arabic Speech Recognition System17 Oct 2023 0 repositories listed
-
Detecting Speech Abnormalities with a Perceiver-based Sequence Classifier that Leverages a Universal Speech Model16 Oct 2023 0 repositories listed
-
End-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis16 Oct 2023 0 repositories listed
-
Personalization of CTC-based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization16 Oct 2023 0 repositories listed
-
Large Vocabulary Spontaneous Speech Recognition for Tigrigna15 Oct 2023 0 repositories listed
-
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring14 Oct 2023 0 repositories listed
-
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text12 Oct 2023 0 repositories listed
-
On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition12 Oct 2023 0 repositories listed
-
Acoustic Model Fusion for End-to-end Speech Recognition10 Oct 2023 0 repositories listed
-
Discriminative Speech Recognition Rescoring with Pre-trained Language Models10 Oct 2023 0 repositories listed
-
Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent Synthesis9 Oct 2023 0 repositories listed
-
Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech Recognition7 Oct 2023 0 repositories listed
-
A privacy-preserving method using secret key for convolutional neural network-based speech classification6 Oct 2023 0 repositories listed
-
HuBERTopic: Enhancing Semantic Representation of HuBERT through Self-supervision Utilizing Topic Model6 Oct 2023 0 repositories listed
-
An Integrated Algorithm for Robust and Imperceptible Audio Adversarial Examples5 Oct 2023 0 repositories listed
-
Neural Language Model Pruning for Automatic Speech Recognition5 Oct 2023 0 repositories listed
-
4 Oct 2023 0 repositories listed
-
One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition2 Oct 2023 0 repositories listed
-
AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR30 Sep 2023 0 repositories listed
-
AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition29 Sep 2023 0 repositories listed
-
Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm29 Sep 2023 0 repositories listed
-
Enhancing Code-switching Speech Recognition with Interactive Language Biases29 Sep 2023 0 repositories listed
-
29 Sep 2023 0 repositories listed Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition29 Sep 2023 0 repositories listed
-
The Gift of Feedback: Improving ASR Model Quality by Learning from User Corrections through Federated Learning29 Sep 2023 0 repositories listed
-
Wiki-En-ASR-Adapt: Large-scale synthetic dataset for English ASR Customization29 Sep 2023 0 repositories listed
-
Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR28 Sep 2023 0 repositories listed
-
PP-MeT: a Real-world Personalized Prompt based Meeting Transcription System28 Sep 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.