Browse State-of-the-Art › Speech Recognition › Papers, page 24
Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 6,433 · with a code link: 1,373 · where Syntology ran a sample: 196 (162 with a run with no instrument failure, 34 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (196 of 6,433 tagged: 162 with a run with no instrument failure, 34 where every run was a failure of Syntology's instrument)
Page 24 of 65: papers 2,301 to 2,400 of 6,433, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent Synthesis9 Oct 2023 0 repositories listed
-
End-to-End Lip Reading in Romanian with Cross-Lingual Domain Adaptation and Lateral Inhibition7 Oct 2023 0 repositories listed
-
Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech Recognition7 Oct 2023 0 repositories listed
-
A privacy-preserving method using secret key for convolutional neural network-based speech classification6 Oct 2023 0 repositories listed
-
HuBERTopic: Enhancing Semantic Representation of HuBERT through Self-supervision Utilizing Topic Model6 Oct 2023 0 repositories listed
-
An Integrated Algorithm for Robust and Imperceptible Audio Adversarial Examples5 Oct 2023 0 repositories listed
-
Challenges and Insights: Exploring 3D Spatial Features and Complex Networks on the MISP Dataset5 Oct 2023 0 repositories listed
-
Neural Language Model Pruning for Automatic Speech Recognition5 Oct 2023 0 repositories listed
-
The North System for Formosa Speech Recognition Challenge 20235 Oct 2023 0 repositories listed
-
4 Oct 2023 0 repositories listed
-
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers3 Oct 2023 0 repositories listed
-
One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition2 Oct 2023 0 repositories listed
-
Active Learning Based Fine-Tuning Framework for Speech Emotion Recognition30 Sep 2023 0 repositories listed
-
AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR30 Sep 2023 0 repositories listed
-
SLM: Bridge the thin gap between speech and text foundation models30 Sep 2023 0 repositories listed
-
AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition29 Sep 2023 0 repositories listed
-
Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm29 Sep 2023 0 repositories listed
-
Enhancing Code-switching Speech Recognition with Interactive Language Biases29 Sep 2023 0 repositories listed
-
29 Sep 2023 0 repositories listed Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition29 Sep 2023 0 repositories listed
-
The Gift of Feedback: Improving ASR Model Quality by Learning from User Corrections through Federated Learning29 Sep 2023 0 repositories listed
-
Wiki-En-ASR-Adapt: Large-scale synthetic dataset for English ASR Customization29 Sep 2023 0 repositories listed
-
Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR28 Sep 2023 0 repositories listed
-
PP-MeT: a Real-world Personalized Prompt based Meeting Transcription System28 Sep 2023 0 repositories listed
-
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study27 Sep 2023 0 repositories listed
-
Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study27 Sep 2023 0 repositories listed
-
27 Sep 2023 0 repositories listed
-
Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition26 Sep 2023 0 repositories listed
-
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference26 Sep 2023 0 repositories listed
-
Unsupervised Pre-Training for Vietnamese Automatic Speech Recognition in the HYKIST Project26 Sep 2023 0 repositories listed
-
AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data25 Sep 2023 0 repositories listed
-
Connecting Speech Encoder and Large Language Model for ASR25 Sep 2023 0 repositories listed
-
On the Impact of Quantization and Pruning of Self-Supervised Speech Models for Downstream Speech Recognition Tasks "In-the-Wild''25 Sep 2023 0 repositories listed
-
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers25 Sep 2023 0 repositories listed
-
Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units25 Sep 2023 0 repositories listed
-
Cross-modal Alignment with Optimal Transport for CTC-based ASR24 Sep 2023 0 repositories listed
-
Speech enhancement with frequency domain auto-regressive modeling24 Sep 2023 0 repositories listed
-
My Science Tutor (MyST) -- A Large Corpus of Children's Conversational Speech23 Sep 2023 0 repositories listed
-
Affect Recognition in Conversations Using Large Language Models22 Sep 2023 0 repositories listed
-
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model22 Sep 2023 0 repositories listed
-
Importance of Smoothness Induced by Optimizers in FL4ASR: Towards Understanding Federated Learning for End-to-End ASR22 Sep 2023 0 repositories listed
-
Massive End-to-end Models for Short Search Queries22 Sep 2023 0 repositories listed
-
NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization22 Sep 2023 0 repositories listed
-
A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech Enhancement21 Sep 2023 0 repositories listed
-
Sparsely Shared LoRA on Whisper for Child Speech Recognition21 Sep 2023 0 repositories listed
-
Variational Connectionist Temporal Classification for Order-Preserving Sequence Modeling21 Sep 2023 0 repositories listed
-
AudioFool: Fast, Universal and synchronization-free Cross-Domain Attack on Speech Recognition20 Sep 2023 0 repositories listed
-
Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition20 Sep 2023 0 repositories listed
-
Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition19 Sep 2023 0 repositories listed
-
End-to-End Speech Recognition Contextualization with Large Language Models19 Sep 2023 0 repositories listed
-
Exploring Speech Enhancement for Low-resource Speech Synthesis19 Sep 2023 0 repositories listed
-
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement19 Sep 2023 0 repositories listed
-
Semi-Autoregressive Streaming ASR With Label Context19 Sep 2023 0 repositories listed
-
A Multitask Training Approach to Enhance Whisper with Contextual Biasing and Open-Vocabulary Keyword Spotting18 Sep 2023 0 repositories listed
-
Corpus Synthesis for Zero-shot ASR domain Adaptation using Large Language Models18 Sep 2023 0 repositories listed
-
Distilling HuBERT with LSTMs via Decoupled Knowledge Distillation18 Sep 2023 0 repositories listed
-
Enhancing Multilingual Speech Recognition through Language Prompt Tuning and Frame-Level Language Adapter18 Sep 2023 0 repositories listed
-
HTEC: Human Transcription Error Correction18 Sep 2023 0 repositories listed
-
Instruction-Following Speech Recognition18 Sep 2023 0 repositories listed
-
Investigating End-to-End ASR Architectures for Long Form Audio Transcription18 Sep 2023 0 repositories listed
-
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning17 Sep 2023 0 repositories listed
-
Boosting End-to-End Multilingual Phoneme Recognition through Exploiting Universal Speech Attributes Constraints16 Sep 2023 0 repositories listed
-
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation16 Sep 2023 0 repositories listed
-
Improving Speech Recognition for African American English With Audio Classification16 Sep 2023 0 repositories listed
-
Augmenting conformers with structured state-space sequence models for online speech recognition15 Sep 2023 0 repositories listed
-
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition15 Sep 2023 0 repositories listed
-
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription15 Sep 2023 0 repositories listed
-
t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability15 Sep 2023 0 repositories listed
-
The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction15 Sep 2023 0 repositories listed
-
Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network15 Sep 2023 0 repositories listed
-
CoLLD: Contrastive Layer-to-layer Distillation for Compressing Multilingual Pre-trained Speech Encoders14 Sep 2023 0 repositories listed
-
CPPF: A contextual and post-processing-free model for automatic speech recognition14 Sep 2023 0 repositories listed
-
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks14 Sep 2023 0 repositories listed
-
Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition14 Sep 2023 0 repositories listed
-
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation14 Sep 2023 0 repositories listed
-
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer14 Sep 2023 0 repositories listed
-
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks14 Sep 2023 0 repositories listed
-
Can Whisper perform speech-based in-context learning?13 Sep 2023 0 repositories listed
-
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis13 Sep 2023 0 repositories listed
-
Open-vocabulary Keyword-spotting with Adaptive Instance Normalization13 Sep 2023 0 repositories listed
-
Co-learning synaptic delays, weights and adaptation in spiking neural networks12 Sep 2023 0 repositories listed
-
Improving Robustness of Neural Inverse Text Normalization via Data-Augmentation, Semi-Supervised Learning, and Post-Aligning Method12 Sep 2023 0 repositories listed
-
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults12 Sep 2023 0 repositories listed
-
Minuteman: Machine and Human Joining Forces in Meeting Summarization11 Sep 2023 0 repositories listed
-
Leveraging Large Language Models for Exploiting ASR Uncertainty9 Sep 2023 0 repositories listed
-
LanSER: Language-Model Supported Speech Emotion Recognition7 Sep 2023 0 repositories listed
-
Multiple Representation Transfer from Large Language Models to End-to-End ASR Systems7 Sep 2023 0 repositories listed
-
RoDia: A New Dataset for Romanian Dialect Identification from Speech6 Sep 2023 0 repositories listed
-
Self-Supervised Masked Digital Elevation Models Encoding for Low-Resource Downstream Tasks6 Sep 2023 0 repositories listed
-
Bring the Noise: Introducing Noise Robustness to Pretrained Automatic Speech Recognition5 Sep 2023 0 repositories listed
-
TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-device ASR Models5 Sep 2023 0 repositories listed
-
AVATAR: Robust Voice Search Engine Leveraging Autoregressive Document Retrieval and Contrastive Learning4 Sep 2023 0 repositories listed
-
SememeASR: Boosting Performance of End-to-End Speech Recognition against Domain and Long-Tailed Data Shift with Sememe Semantic Knowledge4 Sep 2023 0 repositories listed
-
Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation4 Sep 2023 0 repositories listed
-
Mapping AI Arguments in Journalism Studies3 Sep 2023 0 repositories listed
-
Contextual Biasing of Named-Entities with Large Language Models1 Sep 2023 0 repositories listed
-
Learning Speech Representation From Contrastive Token-Acoustic Pretraining1 Sep 2023 0 repositories listed
-
Mi-Go: Test Framework which uses YouTube as Data Source for Evaluating Speech Recognition Models like OpenAI's Whisper1 Sep 2023 0 repositories listed
-
Knowledge Distillation from Non-streaming to Streaming ASR Encoder using Auxiliary Non-streaming Layer31 Aug 2023 0 repositories listed
-
ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers30 Aug 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.