Browse State-of-the-Art › speech-recognition › Papers, page 21
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 21 of 58: papers 2,001 to 2,100 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Artificial Intelligence for Cochlear Implants: Review of Strategies, Challenges, and Perspectives17 Mar 2024 0 repositories listed
-
Energy-Based Models with Applications to Speech and Language Processing16 Mar 2024 0 repositories listed
-
Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR16 Mar 2024 0 repositories listed
-
Hearing-Loss Compensation Using Deep Neural Networks: A Framework and Results From a Listening Test15 Mar 2024 0 repositories listed
-
More than words: Advancements and challenges in speech recognition for singing14 Mar 2024 0 repositories listed
-
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer14 Mar 2024 0 repositories listed
-
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children13 Mar 2024 0 repositories listed
-
Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition13 Mar 2024 0 repositories listed
-
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets12 Mar 2024 0 repositories listed
-
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language12 Mar 2024 0 repositories listed
-
The evaluation of a code-switched Sepedi-English automatic speech recognition system11 Mar 2024 0 repositories listed
-
Aligning Speech to Languages to Enhance Code-switching Speech Recognition9 Mar 2024 0 repositories listed
-
A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain7 Mar 2024 0 repositories listed
-
Classist Tools: Social Class Correlates with Performance in NLP7 Mar 2024 0 repositories listed
-
Non-verbal information in spontaneous speech -- towards a new framework of analysis6 Mar 2024 0 repositories listed
-
RADIA -- Radio Advertisement Detection with Intelligent Analytics6 Mar 2024 0 repositories listed
-
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models5 Mar 2024 0 repositories listed
-
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition4 Mar 2024 0 repositories listed
-
What has LeBenchmark Learnt about French Syntax?4 Mar 2024 0 repositories listed
-
A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement3 Mar 2024 0 repositories listed
-
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey2 Mar 2024 0 repositories listed
-
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview1 Mar 2024 0 repositories listed
-
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition29 Feb 2024 0 repositories listed
-
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems29 Feb 2024 0 repositories listed
-
Exploration of Adapter for Noise Robust Automatic Speech Recognition28 Feb 2024 0 repositories listed
-
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement27 Feb 2024 0 repositories listed
-
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models27 Feb 2024 0 repositories listed
-
24 Feb 2024 0 repositories listed
-
Efficient data selection employing Semantic Similarity-based Graph Structures for model training22 Feb 2024 0 repositories listed
-
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR21 Feb 2024 0 repositories listed
-
Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition20 Feb 2024 0 repositories listed
-
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions20 Feb 2024 0 repositories listed
-
Ain't Misbehavin' -- Using LLMs to Generate Expressive Robot Behavior in Conversations with the Tabletop Robot Haru18 Feb 2024 0 repositories listed
-
Cross-Attention Fusion of Visual and Geometric Features for Large Vocabulary Arabic Lipreading18 Feb 2024 0 repositories listed
-
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives14 Feb 2024 0 repositories listed
-
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models14 Feb 2024 0 repositories listed
-
Syllable based DNN-HMM Cantonese Speech to Text System13 Feb 2024 0 repositories listed
-
SALAD: Smart AI Language Assistant Daily12 Feb 2024 0 repositories listed
-
The Balancing Act: Unmasking and Alleviating ASR Biases in Portuguese12 Feb 2024 0 repositories listed
-
The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models12 Feb 2024 0 repositories listed
-
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition10 Feb 2024 0 repositories listed
-
Self-consistent context aware conformer transducer for speech recognition9 Feb 2024 0 repositories listed
-
Progressive unsupervised domain adaptation for ASR using ensemble models and multi-stage training7 Feb 2024 0 repositories listed
-
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems5 Feb 2024 0 repositories listed
-
Resolving Transcription Ambiguity in Spanish: A Hybrid Acoustic-Lexical System for Punctuation Restoration5 Feb 2024 0 repositories listed
-
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens3 Feb 2024 0 repositories listed
-
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents2 Feb 2024 0 repositories listed
-
Digits micro-model for accurate and secure transactions2 Feb 2024 0 repositories listed
-
Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges2 Feb 2024 0 repositories listed
-
Introduction to speech recognition1 Feb 2024 0 repositories listed
-
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases1 Feb 2024 0 repositories listed
-
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition31 Jan 2024 0 repositories listed
-
Exploring the limits of decoder-only models trained on public speech recognition corpora31 Jan 2024 0 repositories listed
-
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition31 Jan 2024 0 repositories listed
-
Byte Pair Encoding Is All You Need For Automatic Bengali Speech Recognition28 Jan 2024 0 repositories listed
-
Comparison of parameters of vowel sounds of russian and english languages26 Jan 2024 0 repositories listed
-
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline26 Jan 2024 0 repositories listed
-
CNN architecture extraction on edge GPU24 Jan 2024 0 repositories listed
-
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction24 Jan 2024 0 repositories listed
-
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering24 Jan 2024 0 repositories listed
-
Locality enhanced dynamic biasing and sampling strategies for contextual ASR23 Jan 2024 0 repositories listed
-
Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study23 Jan 2024 0 repositories listed
-
Consistency Based Unsupervised Self-training For ASR Personalisation22 Jan 2024 0 repositories listed
-
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers22 Jan 2024 0 repositories listed
-
Using Large Language Model for End-to-End Chinese ASR and NER21 Jan 2024 0 repositories listed
-
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search19 Jan 2024 0 repositories listed
-
Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition19 Jan 2024 0 repositories listed
-
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition18 Jan 2024 0 repositories listed
-
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks18 Jan 2024 0 repositories listed
-
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition18 Jan 2024 0 repositories listed
-
On Speech Pre-emphasis as a Simple and Inexpensive Method to Boost Speech Enhancement17 Jan 2024 0 repositories listed
-
Two-pass Endpoint Detection for Speech Recognition17 Jan 2024 0 repositories listed
-
Improving ASR Contextual Biasing with Guided Attention16 Jan 2024 0 repositories listed
-
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization16 Jan 2024 0 repositories listed
-
NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription16 Jan 2024 0 repositories listed
-
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective16 Jan 2024 0 repositories listed
-
SeMaScore : a new evaluation metric for automatic speech recognition tasks15 Jan 2024 0 repositories listed
-
Promptformer: Prompted Conformer Transducer for ASR14 Jan 2024 0 repositories listed
-
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization13 Jan 2024 0 repositories listed
-
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints12 Jan 2024 0 repositories listed
-
Transcending Controlled Environments Assessing the Transferability of ASRRobust NLU Models to Real-World Applications12 Jan 2024 0 repositories listed
-
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese12 Jan 2024 0 repositories listed
-
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec211 Jan 2024 0 repositories listed
-
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction11 Jan 2024 0 repositories listed
-
Useful Blunders: Can Automated Speech Recognition Errors Improve Downstream Dementia Classification?10 Jan 2024 0 repositories listed
-
Continuously Learning New Words in Automatic Speech Recognition9 Jan 2024 0 repositories listed
-
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators8 Jan 2024 0 repositories listed
-
Exploratory Evaluation of Speech Content Masking8 Jan 2024 0 repositories listed
-
High-precision Voice Search Query Correction via Retrievable Speech-text Embedings8 Jan 2024 0 repositories listed
-
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR8 Jan 2024 0 repositories listed
-
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge7 Jan 2024 0 repositories listed
-
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition7 Jan 2024 0 repositories listed
-
Part-of-Speech Tagger for Bodo Language using Deep Learning approach6 Jan 2024 0 repositories listed
-
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model5 Jan 2024 0 repositories listed
-
Nonlinear functional regression by functional deep neural network with kernel embedding5 Jan 2024 0 repositories listed
-
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks5 Jan 2024 0 repositories listed
-
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition4 Jan 2024 0 repositories listed
-
Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models3 Jan 2024 0 repositories listed
-
The Art of Deception: Robust Backdoor Attack using Dynamic Stacking of Triggers3 Jan 2024 0 repositories listed
-
The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge26 Dec 2023 0 repositories listed