Browse State-of-the-Art › speech-recognition › Papers, page 19
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 19 of 58: papers 1,801 to 1,900 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification19 Jul 2024 0 repositories listed
-
Reexamining Racial Disparities in Automatic Speech Recognition Performance: The Role of Confounding by Provenance19 Jul 2024 0 repositories listed
-
Handling Numeric Expressions in Automatic Speech Recognition18 Jul 2024 0 repositories listed
-
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR18 Jul 2024 0 repositories listed
-
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training18 Jul 2024 0 repositories listed
-
Robust ASR Error Correction with Conservative Data Filtering18 Jul 2024 0 repositories listed
-
Morphosyntactic Analysis for CHILDES17 Jul 2024 0 repositories listed
-
Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models16 Jul 2024 0 repositories listed
-
The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation16 Jul 2024 0 repositories listed
-
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data15 Jul 2024 0 repositories listed
-
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation14 Jul 2024 0 repositories listed
-
Tamil Language Computing: the Present and the Future11 Jul 2024 0 repositories listed
-
Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition10 Jul 2024 0 repositories listed
-
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks10 Jul 2024 0 repositories listed
-
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing10 Jul 2024 0 repositories listed
-
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation8 Jul 2024 0 repositories listed
-
Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation8 Jul 2024 0 repositories listed
-
Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments7 Jul 2024 0 repositories listed
-
LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech5 Jul 2024 0 repositories listed
-
Multitaper mel-spectrograms for keyword spotting5 Jul 2024 0 repositories listed
-
Romanization Encoding For Multilingual ASR5 Jul 2024 0 repositories listed
-
5 Jul 2024 0 repositories listed
-
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter5 Jul 2024 0 repositories listed
-
Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models5 Jul 2024 0 repositories listed
-
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models5 Jul 2024 0 repositories listed
-
Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation4 Jul 2024 0 repositories listed
-
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis4 Jul 2024 0 repositories listed
-
Serialized Output Training by Learned Dominance4 Jul 2024 0 repositories listed
-
Advanced Framework for Animal Sound Classification With Features Optimization3 Jul 2024 0 repositories listed
-
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations3 Jul 2024 0 repositories listed
-
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition3 Jul 2024 0 repositories listed
-
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition3 Jul 2024 0 repositories listed
-
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge2 Jul 2024 0 repositories listed
-
Towards the Next Frontier in Speech Representation Learning Using Disentanglement2 Jul 2024 0 repositories listed
-
Cross-Lingual Transfer Learning for Speech Translation1 Jul 2024 0 repositories listed
-
Toward Automated Detection of Biased Social Signals from the Content of Clinical Conversations1 Jul 2024 0 repositories listed
-
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations30 Jun 2024 0 repositories listed
-
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition29 Jun 2024 0 repositories listed
-
Open-Source Conversational AI with SpeechBrain 1.029 Jun 2024 0 repositories listed
-
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data28 Jun 2024 0 repositories listed
-
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over27 Jun 2024 0 repositories listed
-
Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment27 Jun 2024 0 repositories listed
-
Automatic Speech Recognition for Hindi26 Jun 2024 0 repositories listed
-
Dynamic Data Pruning for Automatic Speech Recognition26 Jun 2024 0 repositories listed
-
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research26 Jun 2024 0 repositories listed
-
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR26 Jun 2024 0 repositories listed
-
A Comprehensive Solution to Connect Speech Encoder and Large Language Model for ASR25 Jun 2024 0 repositories listed
-
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization25 Jun 2024 0 repositories listed
-
Sequential Editing for Lifelong Training of Speech Recognition Models25 Jun 2024 0 repositories listed
-
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 202424 Jun 2024 0 repositories listed
-
Investigating Confidence Estimation Measures for Speaker Diarization24 Jun 2024 0 repositories listed
-
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss23 Jun 2024 0 repositories listed
-
Decoder-only Architecture for Streaming End-to-end Speech Recognition23 Jun 2024 0 repositories listed
-
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment22 Jun 2024 0 repositories listed
-
InterBiasing: Boost Unseen Word Recognition through Biasing Intermediate Predictions21 Jun 2024 0 repositories listed
-
Perception of Phonological Assimilation by Neural Speech Recognition Models21 Jun 2024 0 repositories listed
-
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices21 Jun 2024 0 repositories listed
-
An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks20 Jun 2024 0 repositories listed
-
DASB -- Discrete Audio and Speech Benchmark20 Jun 2024 0 repositories listed
-
Intelligent Interface: Enhancing Lecture Engagement with Didactic Activity Summaries20 Jun 2024 0 repositories listed
-
Children's Speech Recognition through Discrete Token Enhancement19 Jun 2024 0 repositories listed
-
Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control19 Jun 2024 0 repositories listed
-
ManWav: The First Manchu ASR Model19 Jun 2024 0 repositories listed
-
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model18 Jun 2024 0 repositories listed
-
Performant ASR Models for Medical Entities in Accented Speech18 Jun 2024 0 repositories listed
-
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting18 Jun 2024 0 repositories listed
-
Transcribe, Align and Segment: Creating speech datasets for low-resource languages18 Jun 2024 0 repositories listed
-
Automatic Speech Recognition for Biomedical Data in Bengali Language16 Jun 2024 0 repositories listed
-
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving16 Jun 2024 0 repositories listed
-
Imperceptible Rhythm Backdoor Attacks: Exploring Rhythm Transformation for Embedding Undetectable Vulnerabilities on Speech Recognition16 Jun 2024 0 repositories listed
-
Large Language Models for Dysfluency Detection in Stuttered Speech16 Jun 2024 0 repositories listed
-
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare15 Jun 2024 0 repositories listed
-
Trading Devil: Robust backdoor attack via Stochastic investment models and Bayesian approach15 Jun 2024 0 repositories listed
-
An efficient text augmentation approach for contextualized Mandarin speech recognition14 Jun 2024 0 repositories listed
-
CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge14 Jun 2024 0 repositories listed
-
Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation14 Jun 2024 0 repositories listed
-
Learning Language Structures through Grounding14 Jun 2024 0 repositories listed
-
On the Evaluation of Speech Foundation Models for Spoken Language Understanding14 Jun 2024 0 repositories listed
-
Optimizing Byte-level Representation for End-to-end ASR14 Jun 2024 0 repositories listed
-
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition14 Jun 2024 0 repositories listed
-
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR14 Jun 2024 0 repositories listed
-
AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers13 Jun 2024 0 repositories listed
-
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech13 Jun 2024 0 repositories listed
-
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment13 Jun 2024 0 repositories listed
-
Multi-Modal Retrieval For Large Language Model Based Speech Recognition13 Jun 2024 0 repositories listed
-
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time13 Jun 2024 0 repositories listed
-
The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments13 Jun 2024 0 repositories listed
-
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition13 Jun 2024 0 repositories listed
-
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data12 Jun 2024 0 repositories listed
-
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness12 Jun 2024 0 repositories listed
-
Dual-Pipeline with Low-Rank Adaptation for New Language Integration in Multilingual ASR12 Jun 2024 0 repositories listed
-
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion12 Jun 2024 0 repositories listed
-
Improving child speech recognition with augmented child-like speech12 Jun 2024 0 repositories listed
-
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets12 Jun 2024 0 repositories listed
-
Neural Blind Source Separation and Diarization for Distant Speech Recognition12 Jun 2024 0 repositories listed
-
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models12 Jun 2024 0 repositories listed
-
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding12 Jun 2024 0 repositories listed
-
Refining Self-Supervised Learnt Speech Representation using Brain Activations12 Jun 2024 0 repositories listed
-
Transformer-based Model for ASR N-Best Rescoring and Rewriting12 Jun 2024 0 repositories listed
-
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection11 Jun 2024 0 repositories listed