Browse State-of-the-Art › speech-recognition › Papers, page 17
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 17 of 58: papers 1,601 to 1,700 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge21 Nov 2024 0 repositories listed
-
CAFE A Novel Code switching Dataset for Algerian Dialect French and English20 Nov 2024 0 repositories listed
-
From Statistical Methods to Pre-Trained Models; A Survey on Automatic Speech Recognition for Resource Scarce Urdu Language20 Nov 2024 0 repositories listed
-
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM20 Nov 2024 0 repositories listed
-
Towards Advanced Speech Signal Processing: A Statistical Perspective on Convolution-Based Architectures and its Applications20 Nov 2024 0 repositories listed
-
Whisper Finetuning on Nepali Language19 Nov 2024 0 repositories listed
-
A Novel Speech Analysis and Correction Tool for Arabic-Speaking Children18 Nov 2024 0 repositories listed
-
Inter-linguistic Phonetic Composition (IPC): A Theoretical and Computational Approach to Enhance Second Language Pronunciation17 Nov 2024 0 repositories listed
-
DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization15 Nov 2024 0 repositories listed
-
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems15 Nov 2024 0 repositories listed
-
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data14 Nov 2024 0 repositories listed
-
Transferable Adversarial Attacks against ASR14 Nov 2024 0 repositories listed
-
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions11 Nov 2024 0 repositories listed
-
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages7 Nov 2024 0 repositories listed
-
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO1 Nov 2024 0 repositories listed
-
Optimizing Contextual Speech Recognition Using Vector Quantization for Efficient Retrieval1 Nov 2024 0 repositories listed
-
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?31 Oct 2024 0 repositories listed
-
Augmenting Polish Automatic Speech Recognition System With Synthetic Data30 Oct 2024 0 repositories listed
-
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising30 Oct 2024 0 repositories listed
-
Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription29 Oct 2024 0 repositories listed
-
Asynchronous Tool Usage for Real-Time Agents28 Oct 2024 0 repositories listed
-
Multilingual Standalone Trustworthy Voice-Based Social Network for Disaster Situations28 Oct 2024 0 repositories listed
-
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs27 Oct 2024 0 repositories listed
-
A Survey on Speech Large Language Models24 Oct 2024 0 repositories listed
-
Contextual Biasing to Improve Domain-specific Custom Vocabulary Audio Transcription without Explicit Fine-Tuning of Whisper Model24 Oct 2024 0 repositories listed
-
Evaluating and Improving Automatic Speech Recognition Systems for Korean Meteorological Experts24 Oct 2024 0 repositories listed
-
kNN For Whisper And Its Effect On Bias And Speaker Adaptation24 Oct 2024 0 repositories listed
-
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams23 Oct 2024 0 repositories listed
-
DENOASR: Debiasing ASRs through Selective Denoising22 Oct 2024 0 repositories listed
-
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap22 Oct 2024 0 repositories listed
-
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models22 Oct 2024 0 repositories listed
-
Acoustic Model Optimization over Multiple Data Sources: Merging and Valuation21 Oct 2024 0 repositories listed
-
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding21 Oct 2024 0 repositories listed
-
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach19 Oct 2024 0 repositories listed
-
AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup18 Oct 2024 0 repositories listed
-
Computational Approaches to Arabic-English Code-Switching17 Oct 2024 0 repositories listed
-
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation17 Oct 2024 0 repositories listed
-
Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR17 Oct 2024 0 repositories listed
-
Roadmap towards Superhuman Speech Understanding using Large Language Models17 Oct 2024 0 repositories listed
-
Investigation of Speaker Representation for Target-Speaker Speech Processing15 Oct 2024 0 repositories listed
-
Character-aware audio-visual subtitling in context14 Oct 2024 0 repositories listed
-
In-Materia Speech Recognition14 Oct 2024 0 repositories listed
-
State of NLP in Kenya: A Survey13 Oct 2024 0 repositories listed
-
Automatic Speech Recognition with BERT and CTC Transformers: A Review12 Oct 2024 0 repositories listed
-
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities11 Oct 2024 0 repositories listed
-
UniGlyph: A Seven-Segment Script for Universal Language Representation11 Oct 2024 0 repositories listed
-
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models10 Oct 2024 0 repositories listed
-
A two-stage transliteration approach to improve performance of a multilingual ASR9 Oct 2024 0 repositories listed
-
Advocating Character Error Rate for Multilingual ASR Evaluation9 Oct 2024 0 repositories listed
-
The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge8 Oct 2024 0 repositories listed
-
Automatic Screening for Children with Speech Disorder using Automatic Speech Recognition: Opportunities and Challenges7 Oct 2024 0 repositories listed
-
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments7 Oct 2024 0 repositories listed
-
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition6 Oct 2024 0 repositories listed
-
Punctuation Prediction for Polish Texts using Transformers6 Oct 2024 0 repositories listed
-
Enhancement of Dysarthric Speech Reconstruction by Contrastive Learning5 Oct 2024 0 repositories listed
-
The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities5 Oct 2024 0 repositories listed
-
Reverb: Open-Source ASR and Diarization from Rev4 Oct 2024 0 repositories listed
-
Team MTS @ AutoMin 2021: An Overview of Existing Summarization Approaches and Comparison to Unsupervised Summarization Techniques4 Oct 2024 0 repositories listed
-
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings3 Oct 2024 0 repositories listed
-
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems3 Oct 2024 0 repositories listed
-
Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition3 Oct 2024 0 repositories listed
-
HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR3 Oct 2024 0 repositories listed
-
Efficient Streaming LLM for Speech Recognition2 Oct 2024 0 repositories listed
-
Spoken Grammar Assessment Using LLM2 Oct 2024 0 repositories listed
-
Automatic Speech Recognition for the Ika Language1 Oct 2024 0 repositories listed
-
Alignment-Free Training for Transducer-based Multi-Talker ASR30 Sep 2024 0 repositories listed
-
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding30 Sep 2024 0 repositories listed
-
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems30 Sep 2024 0 repositories listed
-
Efficient Long-Form Speech Recognition for General Speech In-Context Learning29 Sep 2024 0 repositories listed
-
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility29 Sep 2024 0 repositories listed
-
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective29 Sep 2024 0 repositories listed
-
Advanced Clustering Techniques for Speech Signal Enhancement: A Review and Metanalysis of Fuzzy C-Means, K-Means, and Kernel Fuzzy C-Means Methods28 Sep 2024 0 repositories listed
-
A GEN AI Framework for Medical Note Generation27 Sep 2024 0 repositories listed
-
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models27 Sep 2024 0 repositories listed
-
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study26 Sep 2024 0 repositories listed
-
Deep CLAS: Deep Contextual Listen, Attend and Spell26 Sep 2024 0 repositories listed
-
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition26 Sep 2024 0 repositories listed
-
Unveiling the Role of Pretraining in Direct Speech Translation26 Sep 2024 0 repositories listed
-
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not25 Sep 2024 0 repositories listed
-
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events25 Sep 2024 0 repositories listed
-
Speech Recognition Rescoring with Large Speech-Text Foundation Models25 Sep 2024 0 repositories listed
-
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM24 Sep 2024 0 repositories listed
-
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs24 Sep 2024 0 repositories listed
-
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens24 Sep 2024 0 repositories listed
-
Revisiting Acoustic Features for Robust ASR24 Sep 2024 0 repositories listed
-
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices24 Sep 2024 0 repositories listed
-
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition21 Sep 2024 0 repositories listed
-
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering20 Sep 2024 0 repositories listed
-
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper20 Sep 2024 0 repositories listed
-
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction20 Sep 2024 0 repositories listed
-
LM-assisted keyword biasing with Aho-Corasick algorithm for Transducer-based ASR20 Sep 2024 0 repositories listed
-
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection20 Sep 2024 0 repositories listed
-
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space19 Sep 2024 0 repositories listed
-
Personalized Speech Recognition for Children with Test-Time Adaptation19 Sep 2024 0 repositories listed
-
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts19 Sep 2024 0 repositories listed
-
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR18 Sep 2024 0 repositories listed
-
A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework17 Sep 2024 0 repositories listed
-
Bio-Inspired Mamba: Temporal Locality and Bioplausible Learning in Selective State Space Models17 Sep 2024 0 repositories listed
-
Chain-of-Thought Prompting for Speech Translation17 Sep 2024 0 repositories listed
-
Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models17 Sep 2024 0 repositories listed