Browse State-of-the-Art › speech-recognition › Papers, page 14
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 14 of 58: papers 1,301 to 1,400 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 202516 Jun 2025 0 repositories listed
-
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems16 Jun 2025 0 repositories listed
-
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models16 Jun 2025 0 repositories listed
-
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition15 Jun 2025 0 repositories listed
-
Enabling automatic transcription of child-centered audio recordings from real-world environments13 Jun 2025 0 repositories listed
-
Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform13 Jun 2025 0 repositories listed
-
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test13 Jun 2025 0 repositories listed
-
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition12 Jun 2025 0 repositories listed
-
Improving Named Entity Transcription with Contextual LLM-based Revision12 Jun 2025 0 repositories listed
-
Joint ASR and Speaker Role Tagging with Serialized Output Training12 Jun 2025 0 repositories listed
-
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary11 Jun 2025 0 repositories listed
-
Regularizing Learnable Feature Extraction for Automatic Speech Recognition11 Jun 2025 0 repositories listed
-
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research10 Jun 2025 0 repositories listed
-
Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech9 Jun 2025 0 repositories listed
-
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition9 Jun 2025 0 repositories listed
-
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation9 Jun 2025 0 repositories listed
-
Uncovering the Functional Roles of Nonlinearity in Memory9 Jun 2025 0 repositories listed
-
Unified Semi-Supervised Pipeline for Automatic Speech Recognition9 Jun 2025 0 repositories listed
-
Speech Recognition on TV Series with Video-guided Post-Correction8 Jun 2025 0 repositories listed
-
Automatic Speech Recognition of African American English: Lexical and Contextual Effects7 Jun 2025 0 repositories listed
-
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs7 Jun 2025 0 repositories listed
-
AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition6 Jun 2025 0 repositories listed
-
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition6 Jun 2025 0 repositories listed
-
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models6 Jun 2025 0 repositories listed
-
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems6 Jun 2025 0 repositories listed
-
Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning6 Jun 2025 0 repositories listed
-
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM5 Jun 2025 0 repositories listed
-
Customizing Speech Recognition Model with Large Language Model Feedback5 Jun 2025 0 repositories listed
-
LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models5 Jun 2025 0 repositories listed
-
LLM-based phoneme-to-grapheme for phoneme-based speech recognition5 Jun 2025 0 repositories listed
-
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition5 Jun 2025 0 repositories listed
-
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR4 Jun 2025 0 repositories listed
-
Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts4 Jun 2025 0 repositories listed
-
MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition4 Jun 2025 0 repositories listed
-
A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation3 Jun 2025 0 repositories listed
-
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss3 Jun 2025 0 repositories listed
-
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning3 Jun 2025 0 repositories listed
-
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation2 Jun 2025 0 repositories listed
-
Cocktail-Party Audio-Visual Speech Recognition2 Jun 2025 0 repositories listed
-
DNCASR: End-to-End Training for Speaker-Attributed ASR2 Jun 2025 0 repositories listed
-
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation2 Jun 2025 0 repositories listed
-
Riemannian Time Warping: Multiple Sequence Alignment in Curved Spaces2 Jun 2025 0 repositories listed
-
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric2 Jun 2025 0 repositories listed
-
TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge2 Jun 2025 0 repositories listed
-
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing2 Jun 2025 0 repositories listed
-
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data2 Jun 2025 0 repositories listed
-
Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody1 Jun 2025 0 repositories listed
-
Causal Structure Discovery for Error Diagnostics of Children's ASR31 May 2025 0 repositories listed
-
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems31 May 2025 0 repositories listed
-
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition31 May 2025 0 repositories listed
-
No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction31 May 2025 0 repositories listed
-
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization30 May 2025 0 repositories listed
-
Fewer Hallucinations, More Verification: A Three-Stage LLM-Based Framework for ASR Error Correction30 May 2025 0 repositories listed
-
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC30 May 2025 0 repositories listed
-
MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition30 May 2025 0 repositories listed
-
MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR30 May 2025 0 repositories listed
-
Running Conventional Automatic Speech Recognition on Memristor Hardware: A Simulated Approach30 May 2025 0 repositories listed
-
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation29 May 2025 0 repositories listed
-
Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection29 May 2025 0 repositories listed
-
Advancing Hearing Assessment: An ASR-Based Frequency-Specific Speech Test for Diagnosing Presbycusis28 May 2025 0 repositories listed
-
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition28 May 2025 0 repositories listed
-
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding28 May 2025 0 repositories listed
-
CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge27 May 2025 0 repositories listed
-
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing27 May 2025 0 repositories listed
-
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis27 May 2025 0 repositories listed
-
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use27 May 2025 0 repositories listed
-
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems27 May 2025 0 repositories listed
-
Topological Deep Learning for Speech Data27 May 2025 0 repositories listed
-
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation27 May 2025 0 repositories listed
-
Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection26 May 2025 0 repositories listed
-
Continuous Learning for Children's ASR: Overcoming Catastrophic Forgetting with Elastic Weight Consolidation and Synaptic Intelligence26 May 2025 0 repositories listed
-
In-context Language Learning for Endangered Languages in Speech Recognition26 May 2025 0 repositories listed
-
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization26 May 2025 0 repositories listed
-
Languages in Multilingual Speech Foundation Models Align Both Phonetically and Semantically26 May 2025 0 repositories listed
-
Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition26 May 2025 0 repositories listed
-
Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy26 May 2025 0 repositories listed
-
Robust fine-tuning of speech recognition models via model merging: application to disordered speech26 May 2025 0 repositories listed
-
The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages26 May 2025 0 repositories listed
-
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper25 May 2025 0 repositories listed
-
Building a Functional Machine Translation Corpus for Kpelle24 May 2025 0 repositories listed
-
StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos24 May 2025 0 repositories listed
-
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition23 May 2025 0 repositories listed
-
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining23 May 2025 0 repositories listed
-
An Effective Training Framework for Light-Weight Automatic Speech Recognition Models22 May 2025 0 repositories listed
-
Large Language Models based ASR Error Correction for Child Conversations22 May 2025 0 repositories listed
-
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding22 May 2025 0 repositories listed
-
From Weak Labels to Strong Results: Utilizing 5,000 Hours of Noisy Classroom Transcripts with Minimal Accurate Data20 May 2025 0 repositories listed
-
HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing20 May 2025 0 repositories listed
-
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English20 May 2025 0 repositories listed
-
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties20 May 2025 0 repositories listed
-
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach20 May 2025 0 repositories listed
-
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition20 May 2025 0 repositories listed
-
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down19 May 2025 0 repositories listed
-
Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR19 May 2025 0 repositories listed
-
Granary: Speech Recognition and Translation Dataset in 25 European Languages19 May 2025 0 repositories listed
-
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 202519 May 2025 0 repositories listed
-
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio16 May 2025 0 repositories listed
-
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors16 May 2025 0 repositories listed
-
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems16 May 2025 0 repositories listed
-
Automatic Speech Recognition for African Low-Resource Languages: Challenges and Future Directions16 May 2025 0 repositories listed