Browse State-of-the-Art › speech-recognition › Papers, page 26
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 26 of 58: papers 2,501 to 2,600 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Evaluating OpenAI's Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person23 May 2023 0 repositories listed
-
Graph Meets LLM: A Novel Approach to Collaborative Filtering for Robust Conversational Understanding23 May 2023 0 repositories listed
-
Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning23 May 2023 0 repositories listed
-
On the Transferability of Whisper-based Representations for "In-the-Wild" Cross-Task Downstream Speech Applications23 May 2023 0 repositories listed
-
Personalized Predictive ASR for Latency Reduction in Voice Assistants23 May 2023 0 repositories listed
-
Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding23 May 2023 0 repositories listed
-
SE-Bridge: Speech Enhancement with Consistent Brownian Bridge23 May 2023 0 repositories listed
-
TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition23 May 2023 0 repositories listed
-
Debiased Automatic Speech Recognition for Dysarthric Speech via Sample Reweighting with Sample Affinity Test22 May 2023 0 repositories listed
-
GNCformer Enhanced Self-attention for Automatic Speech Recognition22 May 2023 0 repositories listed
-
Modular Domain Adaptation for Conformer-Based Streaming ASR22 May 2023 0 repositories listed
-
Text Generation with Speech Synthesis for ASR Data Augmentation22 May 2023 0 repositories listed
-
CASA-ASR: Context-Aware Speaker-Attributed ASR21 May 2023 0 repositories listed
-
Contextualized End-to-End Speech Recognition with Contextual Phrase Prediction Network21 May 2023 0 repositories listed
-
DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting21 May 2023 0 repositories listed
-
Hystoc: Obtaining word confidences for fusion of end-to-end ASR systems21 May 2023 0 repositories listed
-
21 May 2023 0 repositories listed
-
On the Efficacy and Noise-Robustness of Jointly Learned Speech Emotion and Automatic Speech Recognition21 May 2023 0 repositories listed
-
Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction21 May 2023 0 repositories listed
-
VAKTA-SETU: A Speech-to-Speech Machine Translation Service in Select Indic Languages21 May 2023 0 repositories listed
-
Self-supervised representations in speech-based depression detection20 May 2023 0 repositories listed
-
Language-universal phonetic encoder for low-resource speech recognition19 May 2023 0 repositories listed
-
Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition19 May 2023 0 repositories listed
-
Unsupervised ASR via Cross-Lingual Pseudo-Labeling19 May 2023 0 repositories listed
-
A Lexical-aware Non-autoregressive Transformer-based ASR Model18 May 2023 0 repositories listed
-
Accurate and Reliable Confidence Estimation Based on Non-Autoregressive End-to-End Speech Recognition System18 May 2023 0 repositories listed
-
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark18 May 2023 0 repositories listed
-
Use of Speech Impairment Severity for Dysarthric Speech Recognition18 May 2023 0 repositories listed
-
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition18 May 2023 0 repositories listed
-
Boosting Local Spectro-Temporal Features for Speech Analysis17 May 2023 0 repositories listed
-
Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation Based Voice Conversion16 May 2023 0 repositories listed
-
Application-Agnostic Language Modeling for On-Device ASR16 May 2023 0 repositories listed
-
Critical Appraisal of Artificial Intelligence-Mediated Communication15 May 2023 0 repositories listed
-
OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking15 May 2023 0 repositories listed
-
Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations14 May 2023 0 repositories listed
-
Accelerator-Aware Training for Transducer-Based Speech Recognition12 May 2023 0 repositories listed
-
Continual Learning for End-to-End ASR by Averaging Domain Experts12 May 2023 0 repositories listed
-
Investigating the Sensitivity of Automatic Speech Recognition Systems to Phonetic Variation in L2 Englishes12 May 2023 0 repositories listed
-
Masked Audio Text Encoders are Effective Multi-Modal Rescorers11 May 2023 0 repositories listed
-
Quran Recitation Recognition using End-to-End Deep Learning10 May 2023 0 repositories listed
-
Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models9 May 2023 0 repositories listed
-
Robust Acoustic and Semantic Contextual Biasing in Neural Transducers for Speech Recognition9 May 2023 0 repositories listed
-
Who Needs Decoders? Efficient Estimation of Sequence-level Attributes9 May 2023 0 repositories listed
-
8 May 2023 0 repositories listed
-
Multi-Temporal Lip-Audio Memory for Visual Speech Recognition8 May 2023 0 repositories listed
-
Neural Steerer: Novel Steering Vector Synthesis with a Causal Neural Field over Frequency and Source Positions8 May 2023 0 repositories listed
-
Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers7 May 2023 0 repositories listed
-
Employing Hybrid Deep Neural Networks on Dari Speech4 May 2023 0 repositories listed
-
End-to-end spoken language understanding using joint CTC loss and self-supervised, pretrained acoustic encoders4 May 2023 0 repositories listed
-
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks4 May 2023 0 repositories listed
-
Considerations for Ethical Speech Recognition Datasets3 May 2023 0 repositories listed
-
A Study on the Integration of Pipeline and E2E SLU systems for Spoken Semantic Parsing toward STOP Quality Challenge2 May 2023 0 repositories listed
-
Lessons Learned in ATCO2: 5000 hours of Air Traffic Control Communications for Robust Automatic Speech Recognition and Understanding2 May 2023 0 repositories listed
-
A Review of Deep Learning Techniques for Speech Processing30 Apr 2023 0 repositories listed
-
Building a Non-native Speech Corpus Featuring Chinese-English Bilingual Children: Compilation and Rationale30 Apr 2023 0 repositories listed
-
Deep Learning-based Spatio Temporal Facial Feature Visual Speech Recognition30 Apr 2023 0 repositories listed
-
Deep Transfer Learning for Automatic Speech Recognition: Towards Better Generalization27 Apr 2023 0 repositories listed
-
Understanding Shared Speech-Text Representations27 Apr 2023 0 repositories listed
-
Modeling Spoken Information Queries for Virtual Assistants: Open Problems, Challenges and Opportunities25 Apr 2023 0 repositories listed
-
Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition24 Apr 2023 0 repositories listed
-
Non-autoregressive End-to-end Approaches for Joint Automatic Speech Recognition and Spoken Language Understanding21 Apr 2023 0 repositories listed
-
Towards the Universal Defense for Query-Based Audio Adversarial Attacks20 Apr 2023 0 repositories listed
-
Security and Privacy Problems in Voice Assistant Applications: A Survey19 Apr 2023 0 repositories listed
-
Approximate Nearest Neighbour Phrase Mining for Contextual Speech Recognition18 Apr 2023 0 repositories listed
-
Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR18 Apr 2023 0 repositories listed
-
Towards the Transferable Audio Adversarial Attack via Ensemble Methods18 Apr 2023 0 repositories listed
-
Multimodal Short Video Rumor Detection System Based on Contrastive Learning17 Apr 2023 0 repositories listed
-
A Virtual Simulation-Pilot Agent for Training of Air Traffic Controllers16 Apr 2023 0 repositories listed
-
A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition15 Apr 2023 0 repositories listed
-
Evaluation of Speaker Anonymization on Emotional Speech15 Apr 2023 0 repositories listed
-
Task-oriented Document-Grounded Dialog Systems by HLTPR@RWTH for DSTC9 and DSTC1014 Apr 2023 0 repositories listed
-
Solving Tensor Low Cycle Rank Approximation13 Apr 2023 0 repositories listed
-
Regularizing Contrastive Predictive Coding for Speech Applications12 Apr 2023 0 repositories listed
-
Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training12 Apr 2023 0 repositories listed
-
Sim-T: Simplify the Transformer Network by Multiplexing Technique for Speech Recognition11 Apr 2023 0 repositories listed
-
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR11 Apr 2023 0 repositories listed
-
Adaptive Feature Fusion: Enhancing Generalization in Deep Learning Models4 Apr 2023 0 repositories listed
-
Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data4 Apr 2023 0 repositories listed
-
Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition3 Apr 2023 0 repositories listed
-
Self-Supervised Learning-Based Source Separation for Meeting Data3 Apr 2023 0 repositories listed
-
Multilingual Word Error Rate Estimation: e-WER32 Apr 2023 0 repositories listed
-
Dialog act guided contextual adapter for personalized speech recognition31 Mar 2023 0 repositories listed
-
31 Mar 2023 0 repositories listed
-
Lego-Features: Exporting modular encoder features for streaming and deliberation ASR31 Mar 2023 0 repositories listed
-
31 Mar 2023 0 repositories listed
-
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers30 Mar 2023 0 repositories listed
-
SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision30 Mar 2023 0 repositories listed
-
AVFormer: Injecting Vision into Frozen Speech Models for Zero-Shot AV-ASR29 Mar 2023 0 repositories listed
-
Joint unsupervised and supervised learning for context-aware language identification29 Mar 2023 0 repositories listed
-
Text is All You Need: Personalizing ASR Models using Controllable Speech Synthesis27 Mar 2023 0 repositories listed
-
A Deliberation-based Joint Acoustic and Text Decoder23 Mar 2023 0 repositories listed
-
23 Mar 2023 0 repositories listed
-
Enhancing Unsupervised Speech Recognition with Diffusion GANs23 Mar 2023 0 repositories listed
-
MSAT: Biologically Inspired Multi-Stage Adaptive Threshold for Conversion of Spiking Neural Networks23 Mar 2023 0 repositories listed
-
Pyramid Multi-branch Fusion DCNN with Multi-Head Self-Attention for Mandarin Speech Recognition23 Mar 2023 0 repositories listed
-
Exploring Turkish Speech Recognition via Hybrid CTC/Attention Architecture and Multi-feature Fusion Network22 Mar 2023 0 repositories listed
-
Self-supervised Learning with Speech Modulation Dropout22 Mar 2023 0 repositories listed
-
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations21 Mar 2023 0 repositories listed
-
Transformers in Speech Processing: A Survey21 Mar 2023 0 repositories listed
-
Code-Switching Text Generation and Injection in Mandarin-English ASR20 Mar 2023 0 repositories listed