Browse State-of-the-Art › speech-recognition › Papers, page 22
speech-recognition
Papers archive 2025-07-28
archive papers tagged: 5,715 · with a code link: 1,277 · where Syntology ran a sample: 162 (134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (162 of 5,715 tagged: 134 with a run with no instrument failure, 28 where every run was a failure of Syntology's instrument)
Page 22 of 58: papers 2,101 to 2,200 of 5,715, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Towards Probing Contact Center Large Language Models26 Dec 2023 0 repositories listed
-
Exploring data augmentation in bias mitigation against non-native-accented speech24 Dec 2023 0 repositories listed
-
BLSTM-Based Confidence Estimation for End-to-End Speech Recognition22 Dec 2023 0 repositories listed
-
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification22 Dec 2023 0 repositories listed
-
Multi-Sentence Grounding for Long-term Instructional Video21 Dec 2023 0 repositories listed
-
BANSpEmo: A Bangla Emotional Speech Recognition Dataset21 Dec 2023 0 repositories listed
-
Collaborative Learning with Artificial Intelligence Speakers (CLAIS): Pre-Service Elementary Science Teachers' Responses to the Prototype20 Dec 2023 0 repositories listed
-
Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models20 Dec 2023 0 repositories listed
-
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?19 Dec 2023 0 repositories listed
-
SpokesBiz -- an Open Corpus of Conversational Polish19 Dec 2023 0 repositories listed
-
Efficiency-oriented approaches for self-supervised speech representation learning18 Dec 2023 0 repositories listed
-
Generative linguistic representation for spoken language identification18 Dec 2023 0 repositories listed
-
Improved Long-Form Speech Recognition by Jointly Modeling the Primary and Non-primary Speakers18 Dec 2023 0 repositories listed
-
Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition18 Dec 2023 0 repositories listed
-
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices16 Dec 2023 0 repositories listed
-
OAVA: the open audio-visual archives aggregator16 Dec 2023 0 repositories listed
-
Generative Context-aware Fine-tuning of Self-supervised Speech Models15 Dec 2023 0 repositories listed
-
IR-UWB Radar-Based Contactless Silent Speech Recognition of Vowels, Consonants, Words, and Phrases15 Dec 2023 0 repositories listed
-
Leveraging Language ID to Calculate Intermediate CTC Loss for Enhanced Code-Switching Speech Recognition15 Dec 2023 0 repositories listed
-
LiteVSR: Efficient Visual Speech Recognition by Learning from Speech Representations of Unlabeled Data15 Dec 2023 0 repositories listed
-
Phoneme-aware Encoding for Prefix-tree-based Contextual ASR15 Dec 2023 0 repositories listed
-
Attention-Guided Adaptation for Code-Switching Speech Recognition14 Dec 2023 0 repositories listed
-
Audio-visual fine-tuning of audio-only ASR models14 Dec 2023 0 repositories listed
-
FastInject: Injecting Unpaired Text Data into CTC-based ASR training14 Dec 2023 0 repositories listed
-
Towards Automatic Data Augmentation for Disordered Speech Recognition14 Dec 2023 0 repositories listed
-
Efficient Representation of the Activation Space in Deep Neural Networks13 Dec 2023 0 repositories listed
-
On Robustness to Missing Video for Audiovisual Speech Recognition13 Dec 2023 0 repositories listed
-
PhasePerturbation: Speech Data Augmentation via Phase Perturbation for Automatic Speech Recognition13 Dec 2023 0 repositories listed
-
Revisiting the Entropy Semiring for Neural Speech Recognition13 Dec 2023 0 repositories listed
-
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models13 Dec 2023 0 repositories listed
-
Self-supervised Adaptive Pre-training of Multilingual Speech Models for Language and Dialect Identification12 Dec 2023 0 repositories listed
-
The GUA-Speech System Description for CNVSRC Challenge 202312 Dec 2023 0 repositories listed
-
Creating Spoken Dialog Systems in Ultra-Low Resourced Settings11 Dec 2023 0 repositories listed
-
Deep Photonic Reservoir Computer for Speech Recognition11 Dec 2023 0 repositories listed
-
Revisiting the Role of Label Smoothing in Enhanced Text Sentiment Classification11 Dec 2023 0 repositories listed
-
A Review of Hybrid and Ensemble in Deep Learning for Natural Language Processing9 Dec 2023 0 repositories listed
-
Batched Low-Rank Adaptation of Foundation Models9 Dec 2023 0 repositories listed
-
Keyword spotting -- Detecting commands in speech using deep learning9 Dec 2023 0 repositories listed
-
FreqFed: A Frequency Analysis-Based Approach for Mitigating Poisoning Attacks in Federated Learning7 Dec 2023 0 repositories listed
-
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition6 Dec 2023 0 repositories listed
-
Multimodal Data and Resource Efficient Device-Directed Speech Detection with Large Foundation Models6 Dec 2023 0 repositories listed
-
Optimizing Two-Pass Cross-Lingual Transfer Learning: Phoneme Recognition and Phoneme to Grapheme Translation6 Dec 2023 0 repositories listed
-
PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features5 Dec 2023 0 repositories listed
-
End-to-End Speech-to-Text Translation: A Survey2 Dec 2023 0 repositories listed
-
Self Generated Wargame AI: Double Layer Agent Task Planning Based on Large Language Model2 Dec 2023 0 repositories listed
-
Mavericks at NADI 2023 Shared Task: Unravelling Regional Nuances through Dialect Identification using Transformer-based Approach30 Nov 2023 0 repositories listed
-
Adapting OpenAI's Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets29 Nov 2023 0 repositories listed
-
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data29 Nov 2023 0 repositories listed
-
Phonetic-aware speaker embedding for far-field speaker verification27 Nov 2023 0 repositories listed
-
Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching25 Nov 2023 0 repositories listed
-
Weak Alignment Supervision from Hybrid Model Improves End-to-end ASR24 Nov 2023 0 repositories listed
-
Analysis of Visual Features for Continuous Lipreading in Spanish21 Nov 2023 0 repositories listed
-
Soft Random Sampling: A Theoretical and Empirical Analysis21 Nov 2023 0 repositories listed
-
Speaker-Adapted End-to-End Visual Speech Recognition for Continuous Spanish21 Nov 2023 0 repositories listed
-
App for Resume-Based Job Matching with Speech Interviews and Grammar Analysis: A Review20 Nov 2023 0 repositories listed
-
Beyond Boundaries: A Comprehensive Survey of Transferable Attacks on AI Systems20 Nov 2023 0 repositories listed
-
How does end-to-end speech recognition training impact speech enhancement artifacts?20 Nov 2023 0 repositories listed
-
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition19 Nov 2023 0 repositories listed
-
ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding19 Nov 2023 0 repositories listed
-
GhostVec: A New Threat to Speaker Privacy of End-to-End Speech Recognition System17 Nov 2023 0 repositories listed
-
Improving Large-scale Deep Biasing with Phoneme Features and Text-only Data in Streaming Transducer15 Nov 2023 0 repositories listed
-
Multi-channel Conversational Speaker Separation via Neural Diarization15 Nov 2023 0 repositories listed
-
Enhanced Generative Adversarial Networks for Unseen Word Generation from EEG Signals14 Nov 2023 0 repositories listed
-
Retrieve and Copy: Scaling ASR Personalization to Large Catalogs14 Nov 2023 0 repositories listed
-
On the Effectiveness of ASR Representations in Real-world Noisy Speech Emotion Recognition13 Nov 2023 0 repositories listed
-
Towards End-to-End Spoken Grammatical Error Correction9 Nov 2023 0 repositories listed
-
Whisper in Focus: Enhancing Stuttered Speech Classification with Encoder Layer Optimization9 Nov 2023 0 repositories listed
-
1SPU: 1-step Speech Processing Unit8 Nov 2023 0 repositories listed
-
Fine-tuning convergence model in Bengali speech recognition7 Nov 2023 0 repositories listed
-
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning3 Nov 2023 0 repositories listed
-
Server-side Rescoring of Spoken Entity-centric Knowledge Queries for Virtual Assistants2 Nov 2023 0 repositories listed
-
RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios31 Oct 2023 0 repositories listed
-
Combining Language Models For Specialized Domains: A Colorful Approach30 Oct 2023 0 repositories listed
-
MUST: A Multilingual Student-Teacher Learning approach for low-resource speech recognition29 Oct 2023 0 repositories listed
-
Unified Segment-to-Segment Framework for Simultaneous Sequence Generation27 Oct 2023 0 repositories listed
-
Dialect Adaptation and Data Augmentation for Low-Resource ASR: TalTech Systems for the MADASR 2023 Challenge26 Oct 2023 0 repositories listed
-
UniX-Encoder: A Universal X-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing25 Oct 2023 0 repositories listed
-
Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation23 Oct 2023 0 repositories listed
-
Modality Dropout for Multimodal Device Directed Speech Detection using Verbal and Non-Verbal Features23 Oct 2023 0 repositories listed
-
Quantifying the Dialect Gap and its Correlates Across Languages23 Oct 2023 0 repositories listed
-
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation22 Oct 2023 0 repositories listed
-
Intelligibility prediction with a pretrained noise-robust automatic speech recognition model20 Oct 2023 0 repositories listed
-
The CHiME-7 Challenge: System Description and Performance of NeMo Team's DASR System18 Oct 2023 0 repositories listed
-
Unintended Memorization in Large ASR Models, and How to Mitigate It18 Oct 2023 0 repositories listed
-
Advanced accent/dialect identification and accentedness assessment with multi-embedding models and automatic speech recognition17 Oct 2023 0 repositories listed
-
Audio-AdapterFusion: A Task-ID-free Approach for Efficient and Non-Destructive Multi-task Speech Recognition17 Oct 2023 0 repositories listed
-
Correction Focused Language Model Training for Speech Recognition17 Oct 2023 0 repositories listed
-
Generative error correction for code-switching speech recognition using large language models17 Oct 2023 0 repositories listed
-
Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition17 Oct 2023 0 repositories listed
-
Long-form Simultaneous Speech Translation: Thesis Proposal17 Oct 2023 0 repositories listed
-
Multi-stage Large Language Model Correction for Speech Recognition17 Oct 2023 0 repositories listed
-
VoxArabica: A Robust Dialect-Aware Arabic Speech Recognition System17 Oct 2023 0 repositories listed
-
Detecting Speech Abnormalities with a Perceiver-based Sequence Classifier that Leverages a Universal Speech Model16 Oct 2023 0 repositories listed
-
End-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis16 Oct 2023 0 repositories listed
-
Personalization of CTC-based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization16 Oct 2023 0 repositories listed
-
Large Vocabulary Spontaneous Speech Recognition for Tigrigna15 Oct 2023 0 repositories listed
-
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring14 Oct 2023 0 repositories listed
-
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text12 Oct 2023 0 repositories listed
-
On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition12 Oct 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.