Browse State-of-the-Art › Speech Recognition › Papers, page 21
Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 6,433 · with a code link: 1,373 · where Syntology ran a sample: 196 (162 with a run with no instrument failure, 34 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (196 of 6,433 tagged: 162 with a run with no instrument failure, 34 where every run was a failure of Syntology's instrument)
Page 21 of 65: papers 2,001 to 2,100 of 6,433, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
A Parameter-efficient Language Extension Framework for Multilingual ASR10 Jun 2024 0 repositories listed
-
ASTRA: Aligning Speech and Text Representations for Asr without Sampling10 Jun 2024 0 repositories listed
-
Synthetic Query Generation using Large Language Models for Virtual Assistants10 Jun 2024 0 repositories listed
-
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations9 Jun 2024 0 repositories listed
-
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment9 Jun 2024 0 repositories listed
-
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR7 Jun 2024 0 repositories listed
-
Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis7 Jun 2024 0 repositories listed
-
Flexible Multichannel Speech Enhancement for Noise-Robust Frontend6 Jun 2024 0 repositories listed
-
Helsinki Speech Challenge 20246 Jun 2024 0 repositories listed
-
Hypernetworks for Personalizing ASR to Atypical Speech6 Jun 2024 0 repositories listed
-
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores6 Jun 2024 0 repositories listed
-
Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU6 Jun 2024 0 repositories listed
-
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders5 Jun 2024 0 repositories listed
-
Enhancing CTC-based speech recognition with diverse modeling units5 Jun 2024 0 repositories listed
-
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition5 Jun 2024 0 repositories listed
-
Text Injection for Neural Contextual Biasing5 Jun 2024 0 repositories listed
-
Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing4 Jun 2024 0 repositories listed
-
Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping4 Jun 2024 0 repositories listed
-
Keyword-Guided Adaptation of Automatic Speech Recognition4 Jun 2024 0 repositories listed
-
Compute-Efficient Medical Image Classification with Softmax-Free Transformers and Sequence Normalization3 Jun 2024 0 repositories listed
-
Enabling ASR for Low-Resource Languages: A Comprehensive Dataset Creation Approach3 Jun 2024 0 repositories listed
-
YODAS: Youtube-Oriented Dataset for Audio and Speech2 Jun 2024 0 repositories listed
-
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning1 Jun 2024 0 repositories listed
-
Zipper: A Multi-Tower Decoder Architecture for Fusing Modalities29 May 2024 0 repositories listed
-
Augmented Conversation with Embedded Speech-Driven On-the-Fly Referencing in AR28 May 2024 0 repositories listed
-
Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation28 May 2024 0 repositories listed
-
NUTS, NARS, and Speech28 May 2024 0 repositories listed
-
Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition24 May 2024 0 repositories listed
-
Contextualized Automatic Speech Recognition with Dynamic Vocabulary22 May 2024 0 repositories listed
-
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation22 May 2024 0 repositories listed
-
ST-Gait++: Leveraging spatio-temporal convolutions for gait-based emotion recognition on videos22 May 2024 0 repositories listed
-
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish22 May 2024 0 repositories listed
-
Could a Computer Architect Understand our Brain?21 May 2024 0 repositories listed
-
FairLENS: Assessing Fairness in Law Enforcement Speech Recognition21 May 2024 0 repositories listed
-
Non-autoregressive real-time Accent Conversion model with voice cloning21 May 2024 0 repositories listed
-
Continuous Sign Language Recognition with Adapted Conformer via Unsupervised Pretraining20 May 2024 0 repositories listed
-
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models16 May 2024 0 repositories listed
-
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings15 May 2024 0 repositories listed
-
Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer15 May 2024 0 repositories listed
-
Investigating the 'Autoencoder Behavior' in Speech Self-Supervised Models: a focus on HuBERT's Pretraining14 May 2024 0 repositories listed
-
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants14 May 2024 0 repositories listed
-
SpeechVerse: A Large-scale Generalizable Audio Language Model14 May 2024 0 repositories listed
-
Large Language Models for Education: A Survey12 May 2024 0 repositories listed
-
DP-DyLoRA: Fine-Tuning Transformer-Based Models On-Device under Differentially Private Federated Learning using Dynamic Low-Rank Adaptation10 May 2024 0 repositories listed
-
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech10 May 2024 0 repositories listed
-
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition6 May 2024 0 repositories listed
-
Whispy: Adapting STT Whisper Models to Real-Time Environments6 May 2024 0 repositories listed
-
Combining X-Vectors and Bayesian Batch Active Learning: Two-Stage Active Learning Pipeline for Speech Recognition3 May 2024 0 repositories listed
-
Deep Learning Models in Speech Recognition: Measuring GPU Energy Consumption, Impact of Noise and Model Quantization for Edge Deployment2 May 2024 0 repositories listed
-
Efficient Compression of Multitask Multilingual Speech Models2 May 2024 0 repositories listed
-
Improving Membership Inference in ASR Model Auditing with Perturbed Loss Features2 May 2024 0 repositories listed
-
Low-resource speech recognition and dialect identification of Irish in a multi-task framework2 May 2024 0 repositories listed
-
Sequence-to-sequence models in peer-to-peer learning: A practical application2 May 2024 0 repositories listed
-
Active Learning with Task Adaptation Pre-training for Speech Emotion Recognition1 May 2024 0 repositories listed
-
Efficient Sample-Specific Encoder Perturbations1 May 2024 0 repositories listed
-
Does Whisper understand Swiss German? An automatic, qualitative, and human evaluation30 Apr 2024 0 repositories listed
-
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system29 Apr 2024 0 repositories listed
-
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification29 Apr 2024 0 repositories listed
-
Child Speech Recognition in Human-Robot Interaction: Problem Solved?26 Apr 2024 0 repositories listed
-
Automatic Speech Recognition System-Independent Word Error Rate Estimation25 Apr 2024 0 repositories listed
-
Developing Acoustic Models for Automatic Speech Recognition in Swedish25 Apr 2024 0 repositories listed
-
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF25 Apr 2024 0 repositories listed
-
Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices24 Apr 2024 0 repositories listed
-
Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm23 Apr 2024 0 repositories listed
-
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance23 Apr 2024 0 repositories listed
-
Efficient infusion of self-supervised representations in Automatic Speech Recognition19 Apr 2024 0 repositories listed
-
Learn2Talk: 3D Talking Face Learns from 2D Talking Face19 Apr 2024 0 repositories listed
-
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech18 Apr 2024 0 repositories listed
-
Anatomy of Industrial Scale Multilingual ASR15 Apr 2024 0 repositories listed
-
Resilience of Large Language Models for Noisy Instructions15 Apr 2024 0 repositories listed
-
12 Apr 2024 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task12 Apr 2024 0 repositories listed
-
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution11 Apr 2024 0 repositories listed
-
An inclusive review on deep learning techniques and their scope in handwriting recognition10 Apr 2024 0 repositories listed
-
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping10 Apr 2024 0 repositories listed
-
The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge9 Apr 2024 0 repositories listed
-
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition4 Apr 2024 0 repositories listed
-
Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian3 Apr 2024 0 repositories listed
-
Noise Masking Attacks and Defenses for Pretrained Speech Models2 Apr 2024 0 repositories listed
-
Transfer Learning from Whisper for Microscopic Intelligibility Prediction2 Apr 2024 0 repositories listed
-
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models31 Mar 2024 0 repositories listed
-
LV-CTC: Non-autoregressive ASR with CTC and latent variable models28 Mar 2024 0 repositories listed
-
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition28 Mar 2024 0 repositories listed
-
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus27 Mar 2024 0 repositories listed
-
DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech Recognition26 Mar 2024 0 repositories listed
-
Extracting Biomedical Entities from Noisy Audio Transcripts26 Mar 2024 0 repositories listed
-
Grammatical vs Spelling Error Correction: An Investigation into the Responsiveness of Transformer-based Language Models using BART and MarianMT25 Mar 2024 0 repositories listed
-
Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models25 Mar 2024 0 repositories listed
-
Privacy-Preserving End-to-End Spoken Language Understanding22 Mar 2024 0 repositories listed
-
A Multimodal Approach to Device-Directed Speech Detection with Large Language Models21 Mar 2024 0 repositories listed
-
21 Mar 2024 0 repositories listed
-
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception21 Mar 2024 0 repositories listed
-
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech20 Mar 2024 0 repositories listed
-
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning20 Mar 2024 0 repositories listed
-
Open Access NAO (OAN): a ROS2-based software framework for HRI applications with the NAO robot20 Mar 2024 0 repositories listed
-
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition18 Mar 2024 0 repositories listed
-
Artificial Intelligence for Cochlear Implants: Review of Strategies, Challenges, and Perspectives17 Mar 2024 0 repositories listed
-
Energy-Based Models with Applications to Speech and Language Processing16 Mar 2024 0 repositories listed
-
Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR16 Mar 2024 0 repositories listed
-
Hearing-Loss Compensation Using Deep Neural Networks: A Framework and Results From a Listening Test15 Mar 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.