Browse State-of-the-Art › Speech Recognition › Papers, page 28
Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 6,433 · with a code link: 1,373 · where Syntology ran a sample: 196 (162 with a run with no instrument failure, 34 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (196 of 6,433 tagged: 162 with a run with no instrument failure, 34 where every run was a failure of Syntology's instrument)
Page 28 of 65: papers 2,701 to 2,800 of 6,433, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Knowledge Distillation from Multiple Foundation Models for End-to-End Speech Recognition20 Mar 2023 0 repositories listed
-
On-the-fly Text Retrieval for End-to-End ASR Adaptation20 Mar 2023 0 repositories listed
-
A Deep Learning System for Domain-specific Speech Recognition18 Mar 2023 0 repositories listed
-
DistillW2V2: A Small and Streaming Wav2vec 2.0 Based ASR Model16 Mar 2023 0 repositories listed
-
Improving Perceptual Quality, Intelligibility, and Acoustics on VoIP Platforms16 Mar 2023 0 repositories listed
-
Trustera: A Live Conversation Redaction System16 Mar 2023 0 repositories listed
-
Visual Information Matters for ASR Error Correction16 Mar 2023 0 repositories listed
-
A large-scale multimodal dataset of human speech recognition15 Mar 2023 0 repositories listed
-
Sharing Low Rank Conformer Weights for Tiny Always-On Ambient Speech Recognition Models15 Mar 2023 0 repositories listed
-
Dynamic Alignment Mask CTC: Improved Mask-CTC with Aligned Cross Entropy14 Mar 2023 0 repositories listed
-
Improving Accented Speech Recognition with Multi-Domain Training14 Mar 2023 0 repositories listed
-
Context-Aware Selective Label Smoothing for Calibrating Sequence Recognition Model13 Mar 2023 0 repositories listed
-
Improving the Intent Classification accuracy in Noisy Environment12 Mar 2023 0 repositories listed
-
The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge11 Mar 2023 0 repositories listed
-
An Overview on Language Models: Recent Developments and Outlook10 Mar 2023 0 repositories listed
-
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings10 Mar 2023 0 repositories listed
-
MIXPGD: Hybrid Adversarial Training for Speech Recognition Systems10 Mar 2023 0 repositories listed
-
Unsupervised Language agnostic WER Standardization9 Mar 2023 0 repositories listed
-
wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts6 Mar 2023 0 repositories listed
-
End-to-End Speech Recognition: A Survey3 Mar 2023 0 repositories listed
-
Pre-trained Model Representations and their Robustness against Noise for Speech Emotion Analysis3 Mar 2023 0 repositories listed
-
SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks3 Mar 2023 0 repositories listed
-
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages2 Mar 2023 0 repositories listed
-
Leveraging Large Text Corpora for End-to-End Speech Summarization2 Mar 2023 0 repositories listed
-
LiteG2P: A fast, light and high accuracy model for grapheme-to-phoneme conversion2 Mar 2023 0 repositories listed
-
Leveraging Redundancy in Multiple Audio Signals for Far-Field Speech Recognition1 Mar 2023 0 repositories listed
-
N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space1 Mar 2023 0 repositories listed
-
Synthetic Cross-accent Data Augmentation for Automatic Speech Recognition1 Mar 2023 0 repositories listed
-
A Token-Wise Beam Search Algorithm for RNN-T28 Feb 2023 0 repositories listed
-
Exploring Self-supervised Pre-trained ASR Models For Dysarthric and Elderly Speech Recognition28 Feb 2023 0 repositories listed
-
Practice of the conformer enhanced AUDIO-VISUAL HUBERT on Mandarin and English28 Feb 2023 0 repositories listed
-
A Comparison of Speech Data Augmentation Methods Using S3PRL Toolkit27 Feb 2023 0 repositories listed
-
Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video27 Feb 2023 0 repositories listed
-
Diacritic Recognition Performance in Arabic ASR27 Feb 2023 0 repositories listed
-
Diagonal State Space Augmented Transformers for Speech Recognition27 Feb 2023 0 repositories listed
-
Explanations for Automatic Speech Recognition27 Feb 2023 0 repositories listed
-
Improving Medical Speech-to-Text Accuracy with Vision-Language Pre-training Model27 Feb 2023 0 repositories listed
-
MoLE : Mixture of Language Experts for Multi-Lingual Automatic Speech Recognition27 Feb 2023 0 repositories listed
-
From Audio to Symbolic Encoding26 Feb 2023 0 repositories listed
-
Speech Corpora Divergence Based Unsupervised Data Selection for ASR26 Feb 2023 0 repositories listed
-
Chaotic Variational Auto encoder-based Adversarial Machine Learning25 Feb 2023 0 repositories listed
-
Ensemble knowledge distillation of self-supervised speech models24 Feb 2023 0 repositories listed
-
Factual Consistency Oriented Speech Recognition24 Feb 2023 0 repositories listed
-
Evaluating Automatic Speech Recognition in an Incremental Setting23 Feb 2023 0 repositories listed
-
Generalization of Auto-Regressive Hidden Markov Models to Non-Linear Dynamics and Unit Quaternion Observation Space23 Feb 2023 0 repositories listed
-
Improving Contextual Spelling Correction by External Acoustics Attention and Semantic Aware Data Augmentation22 Feb 2023 0 repositories listed
-
MADI: Inter-domain Matching and Intra-domain Discrimination for Cross-domain Speech Recognition22 Feb 2023 0 repositories listed
-
UML: A Universal Monolingual Output Layer for Multilingual ASR22 Feb 2023 0 repositories listed
-
Connecting Humanities and Social Sciences: Applying Language and Speech Technology to Online Panel Surveys21 Feb 2023 0 repositories listed
-
An ASR-free Fluency Scoring Approach with Self-Supervised Learning20 Feb 2023 0 repositories listed
-
Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition20 Feb 2023 0 repositories listed
-
Optimization Methods in Deep Learning: A Comprehensive Overview19 Feb 2023 0 repositories listed
-
Front-End Adapter: Adapting Front-End Input of Speech based Self-Supervised Learning for Speech Recognition18 Feb 2023 0 repositories listed
-
Speaker and Language Change Detection using Wav2vec2 and Whisper18 Feb 2023 0 repositories listed
-
17 Feb 2023 0 repositories listed
-
17 Feb 2023 0 repositories listed
-
Massively Multilingual Shallow Fusion with Large Language Models17 Feb 2023 0 repositories listed
-
Measuring Equality in Machine Learning Security Defenses: A Case Study in Speech Recognition17 Feb 2023 0 repositories listed
-
Adaptable End-to-End ASR Models using Replaceable Internal LMs and Residual Softmax16 Feb 2023 0 repositories listed
-
16 Feb 2023 0 repositories listed
-
JEIT: Joint End-to-End Model and Internal Language Model Training for Speech Recognition16 Feb 2023 0 repositories listed
-
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition16 Feb 2023 0 repositories listed
-
Speaker Change Detection for Transformer Transducer ASR16 Feb 2023 0 repositories listed
-
Stabilising and accelerating light gated recurrent units for automatic speech recognition16 Feb 2023 0 repositories listed
-
ASR Bundestag: A Large-Scale political debate dataset in German12 Feb 2023 0 repositories listed
-
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations10 Feb 2023 0 repositories listed
-
PATCorrect: Non-autoregressive Phoneme-augmented Transformer for ASR Error Correction10 Feb 2023 0 repositories listed
-
Leveraging supplementary text data to kick-start automatic speech recognition system development with limited transcriptions9 Feb 2023 0 repositories listed
-
LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table Lookup7 Feb 2023 0 repositories listed
-
MAC: A unified framework boosting low resource automatic speech recognition5 Feb 2023 0 repositories listed
-
Efficient Domain Adaptation for Speech Foundation Models3 Feb 2023 0 repositories listed
-
Improving Rare Words Recognition through Homophone Extension and Unified Writing for Low-resource Cantonese Speech Recognition2 Feb 2023 0 repositories listed
-
Exploring Attention Map Reuse for Efficient Transformer Neural Networks29 Jan 2023 0 repositories listed
-
Fillers in Spoken Language Understanding: Computational and Psycholinguistic Perspectives25 Jan 2023 0 repositories listed
-
A Comparison of Temporal Encoders for Neuromorphic Keyword Spotting with Few Neurons24 Jan 2023 0 repositories listed
-
A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset21 Jan 2023 0 repositories listed
-
Regeneration Learning: A Learning Paradigm for Data Generation21 Jan 2023 0 repositories listed
-
Language Agnostic Data-Driven Inverse Text Normalization20 Jan 2023 0 repositories listed
-
From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition19 Jan 2023 0 repositories listed
-
Adapting Multilingual Speech Representation Model for a New, Underresourced Language through Multilingual Fine-tuning and Continued Pretraining18 Jan 2023 0 repositories listed
-
BayesSpeech: A Bayesian Transformer Network for Automatic Speech Recognition16 Jan 2023 0 repositories listed
-
Multi-resolution location-based training for multi-channel continuous speech separation16 Jan 2023 0 repositories listed
-
Using Kaldi for Automatic Speech Recognition of Conversational Austrian German16 Jan 2023 0 repositories listed
-
Rationalizing Predictions by Adversarial Information Calibration15 Jan 2023 0 repositories listed
-
Streaming Punctuation: A Novel Punctuation Technique Leveraging Bidirectional Context for Continuous Speech Recognition10 Jan 2023 0 repositories listed
-
FullStop:Punctuation and Segmentation Prediction for Dutch with Transformers9 Jan 2023 0 repositories listed
-
Equivariant and Steerable Neural Networks: A review with special emphasis on the symmetric group8 Jan 2023 0 repositories listed
-
Using External Off-Policy Speech-To-Text Mappings in Contextual End-To-End Automated Speech Recognition6 Jan 2023 0 repositories listed
-
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration1 Jan 2023 0 repositories listed
-
Unsupervised Pre-Training for Vietnamese Automatic Speech Recognition in the HYKIST Project1 Jan 2023 0 repositories listed
-
Sample-Efficient Unsupervised Domain Adaptation of Speech Recognition Systems A case study for Modern Greek31 Dec 2022 0 repositories listed
-
Memory Augmented Lookup Dictionary based Language Modeling for Automatic Speech Recognition30 Dec 2022 0 repositories listed
-
Macro-block dropout for improved regularization in training end-to-end speech recognition models29 Dec 2022 0 repositories listed
-
Don't Be So Sure! Boosting ASR Decoding via Confidence Relaxation27 Dec 2022 0 repositories listed
-
Alignment Entropy Regularization22 Dec 2022 0 repositories listed
-
4D ASR: Joint modeling of CTC, Attention, Transducer, and Mask-Predict decoders21 Dec 2022 0 repositories listed
-
End-to-End Automatic Speech Recognition model for the Sudanese Dialect21 Dec 2022 0 repositories listed
-
21 Dec 2022 0 repositories listed
-
SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding Tasks20 Dec 2022 0 repositories listed
-
Mu²SLAM: Multitask, Multilingual Speech and Language Models19 Dec 2022 0 repositories listed