Browse State-of-the-Art › Automatic Speech Recognition › Papers, page 16
Automatic Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 3,174 · with a code link: 677 · where Syntology ran a sample: 79 (62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (79 of 3,174 tagged: 62 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 16 of 32: papers 1,501 to 1,600 of 3,174, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
End-to-end spoken language understanding using joint CTC loss and self-supervised, pretrained acoustic encoders4 May 2023 0 repositories listed
-
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks4 May 2023 0 repositories listed
-
Considerations for Ethical Speech Recognition Datasets3 May 2023 0 repositories listed
-
A Study on the Integration of Pipeline and E2E SLU systems for Spoken Semantic Parsing toward STOP Quality Challenge2 May 2023 0 repositories listed
-
Lessons Learned in ATCO2: 5000 hours of Air Traffic Control Communications for Robust Automatic Speech Recognition and Understanding2 May 2023 0 repositories listed
-
A Review of Deep Learning Techniques for Speech Processing30 Apr 2023 0 repositories listed
-
Building a Non-native Speech Corpus Featuring Chinese-English Bilingual Children: Compilation and Rationale30 Apr 2023 0 repositories listed
-
Deep Transfer Learning for Automatic Speech Recognition: Towards Better Generalization27 Apr 2023 0 repositories listed
-
Understanding Shared Speech-Text Representations27 Apr 2023 0 repositories listed
-
Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition24 Apr 2023 0 repositories listed
-
Non-autoregressive End-to-end Approaches for Joint Automatic Speech Recognition and Spoken Language Understanding21 Apr 2023 0 repositories listed
-
Towards the Universal Defense for Query-Based Audio Adversarial Attacks20 Apr 2023 0 repositories listed
-
Security and Privacy Problems in Voice Assistant Applications: A Survey19 Apr 2023 0 repositories listed
-
Multimodal Short Video Rumor Detection System Based on Contrastive Learning17 Apr 2023 0 repositories listed
-
A Virtual Simulation-Pilot Agent for Training of Air Traffic Controllers16 Apr 2023 0 repositories listed
-
A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition15 Apr 2023 0 repositories listed
-
Evaluation of Speaker Anonymization on Emotional Speech15 Apr 2023 0 repositories listed
-
Task-oriented Document-Grounded Dialog Systems by HLTPR@RWTH for DSTC9 and DSTC1014 Apr 2023 0 repositories listed
-
Regularizing Contrastive Predictive Coding for Speech Applications12 Apr 2023 0 repositories listed
-
Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training12 Apr 2023 0 repositories listed
-
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR11 Apr 2023 0 repositories listed
-
Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data4 Apr 2023 0 repositories listed
-
Self-Supervised Learning-Based Source Separation for Meeting Data3 Apr 2023 0 repositories listed
-
Multilingual Word Error Rate Estimation: e-WER32 Apr 2023 0 repositories listed
-
Dialog act guided contextual adapter for personalized speech recognition31 Mar 2023 0 repositories listed
-
31 Mar 2023 0 repositories listed
-
31 Mar 2023 0 repositories listed
-
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers30 Mar 2023 0 repositories listed
-
AVFormer: Injecting Vision into Frozen Speech Models for Zero-Shot AV-ASR29 Mar 2023 0 repositories listed
-
Joint unsupervised and supervised learning for context-aware language identification29 Mar 2023 0 repositories listed
-
Text is All You Need: Personalizing ASR Models using Controllable Speech Synthesis27 Mar 2023 0 repositories listed
-
23 Mar 2023 0 repositories listed
-
Enhancing Unsupervised Speech Recognition with Diffusion GANs23 Mar 2023 0 repositories listed
-
Pyramid Multi-branch Fusion DCNN with Multi-Head Self-Attention for Mandarin Speech Recognition23 Mar 2023 0 repositories listed
-
Self-supervised Learning with Speech Modulation Dropout22 Mar 2023 0 repositories listed
-
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations21 Mar 2023 0 repositories listed
-
Transformers in Speech Processing: A Survey21 Mar 2023 0 repositories listed
-
Code-Switching Text Generation and Injection in Mandarin-English ASR20 Mar 2023 0 repositories listed
-
Knowledge Distillation from Multiple Foundation Models for End-to-End Speech Recognition20 Mar 2023 0 repositories listed
-
A Deep Learning System for Domain-specific Speech Recognition18 Mar 2023 0 repositories listed
-
DistillW2V2: A Small and Streaming Wav2vec 2.0 Based ASR Model16 Mar 2023 0 repositories listed
-
Trustera: A Live Conversation Redaction System16 Mar 2023 0 repositories listed
-
Visual Information Matters for ASR Error Correction16 Mar 2023 0 repositories listed
-
Improving Accented Speech Recognition with Multi-Domain Training14 Mar 2023 0 repositories listed
-
Improving the Intent Classification accuracy in Noisy Environment12 Mar 2023 0 repositories listed
-
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings10 Mar 2023 0 repositories listed
-
MIXPGD: Hybrid Adversarial Training for Speech Recognition Systems10 Mar 2023 0 repositories listed
-
wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts6 Mar 2023 0 repositories listed
-
End-to-End Speech Recognition: A Survey3 Mar 2023 0 repositories listed
-
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages2 Mar 2023 0 repositories listed
-
Leveraging Large Text Corpora for End-to-End Speech Summarization2 Mar 2023 0 repositories listed
-
Leveraging Redundancy in Multiple Audio Signals for Far-Field Speech Recognition1 Mar 2023 0 repositories listed
-
N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space1 Mar 2023 0 repositories listed
-
Synthetic Cross-accent Data Augmentation for Automatic Speech Recognition1 Mar 2023 0 repositories listed
-
Practice of the conformer enhanced AUDIO-VISUAL HUBERT on Mandarin and English28 Feb 2023 0 repositories listed
-
A Comparison of Speech Data Augmentation Methods Using S3PRL Toolkit27 Feb 2023 0 repositories listed
-
Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video27 Feb 2023 0 repositories listed
-
Diacritic Recognition Performance in Arabic ASR27 Feb 2023 0 repositories listed
-
Explanations for Automatic Speech Recognition27 Feb 2023 0 repositories listed
-
Improving Medical Speech-to-Text Accuracy with Vision-Language Pre-training Model27 Feb 2023 0 repositories listed
-
MoLE : Mixture of Language Experts for Multi-Lingual Automatic Speech Recognition27 Feb 2023 0 repositories listed
-
Speech Corpora Divergence Based Unsupervised Data Selection for ASR26 Feb 2023 0 repositories listed
-
Ensemble knowledge distillation of self-supervised speech models24 Feb 2023 0 repositories listed
-
Factual Consistency Oriented Speech Recognition24 Feb 2023 0 repositories listed
-
Evaluating Automatic Speech Recognition in an Incremental Setting23 Feb 2023 0 repositories listed
-
Improving Contextual Spelling Correction by External Acoustics Attention and Semantic Aware Data Augmentation22 Feb 2023 0 repositories listed
-
MADI: Inter-domain Matching and Intra-domain Discrimination for Cross-domain Speech Recognition22 Feb 2023 0 repositories listed
-
UML: A Universal Monolingual Output Layer for Multilingual ASR22 Feb 2023 0 repositories listed
-
Connecting Humanities and Social Sciences: Applying Language and Speech Technology to Online Panel Surveys21 Feb 2023 0 repositories listed
-
An ASR-free Fluency Scoring Approach with Self-Supervised Learning20 Feb 2023 0 repositories listed
-
Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition20 Feb 2023 0 repositories listed
-
Speaker and Language Change Detection using Wav2vec2 and Whisper18 Feb 2023 0 repositories listed
-
Massively Multilingual Shallow Fusion with Large Language Models17 Feb 2023 0 repositories listed
-
Adaptable End-to-End ASR Models using Replaceable Internal LMs and Residual Softmax16 Feb 2023 0 repositories listed
-
16 Feb 2023 0 repositories listed
-
Speaker Change Detection for Transformer Transducer ASR16 Feb 2023 0 repositories listed
-
Stabilising and accelerating light gated recurrent units for automatic speech recognition16 Feb 2023 0 repositories listed
-
ASR Bundestag: A Large-Scale political debate dataset in German12 Feb 2023 0 repositories listed
-
PATCorrect: Non-autoregressive Phoneme-augmented Transformer for ASR Error Correction10 Feb 2023 0 repositories listed
-
Leveraging supplementary text data to kick-start automatic speech recognition system development with limited transcriptions9 Feb 2023 0 repositories listed
-
MAC: A unified framework boosting low resource automatic speech recognition5 Feb 2023 0 repositories listed
-
Improving Rare Words Recognition through Homophone Extension and Unified Writing for Low-resource Cantonese Speech Recognition2 Feb 2023 0 repositories listed
-
Fillers in Spoken Language Understanding: Computational and Psycholinguistic Perspectives25 Jan 2023 0 repositories listed
-
A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset21 Jan 2023 0 repositories listed
-
Language Agnostic Data-Driven Inverse Text Normalization20 Jan 2023 0 repositories listed
-
From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition19 Jan 2023 0 repositories listed
-
BayesSpeech: A Bayesian Transformer Network for Automatic Speech Recognition16 Jan 2023 0 repositories listed
-
Multi-resolution location-based training for multi-channel continuous speech separation16 Jan 2023 0 repositories listed
-
Using Kaldi for Automatic Speech Recognition of Conversational Austrian German16 Jan 2023 0 repositories listed
-
Streaming Punctuation: A Novel Punctuation Technique Leveraging Bidirectional Context for Continuous Speech Recognition10 Jan 2023 0 repositories listed
-
Unsupervised Pre-Training for Vietnamese Automatic Speech Recognition in the HYKIST Project1 Jan 2023 0 repositories listed
-
Memory Augmented Lookup Dictionary based Language Modeling for Automatic Speech Recognition30 Dec 2022 0 repositories listed
-
Don't Be So Sure! Boosting ASR Decoding via Confidence Relaxation27 Dec 2022 0 repositories listed
-
Alignment Entropy Regularization22 Dec 2022 0 repositories listed
-
4D ASR: Joint modeling of CTC, Attention, Transducer, and Mask-Predict decoders21 Dec 2022 0 repositories listed
-
End-to-End Automatic Speech Recognition model for the Sudanese Dialect21 Dec 2022 0 repositories listed
-
Mu²SLAM: Multitask, Multilingual Speech and Language Models19 Dec 2022 0 repositories listed
-
Context-aware Fine-tuning of Self-supervised Speech Models16 Dec 2022 0 repositories listed
-
Fast Entropy-Based Methods of Word-Level Confidence Estimation for End-To-End Automatic Speech Recognition16 Dec 2022 0 repositories listed
-
Speech Aware Dialog System Technology Challenge (DSTC11)16 Dec 2022 0 repositories listed