Browse State-of-the-Art › Speech-to-Text › Papers, page 3
Speech-to-Text
Papers archive 2025-07-28
archive papers tagged: 403 · with a code link: 129 · where Syntology ran a sample: 24 (18 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (24 of 403 tagged: 18 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument)
Page 3 of 5: papers 201 to 300 of 403, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Hands-Free VR23 Feb 2024 0 repositories listed
-
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?19 Feb 2024 0 repositories listed
-
Syllable based DNN-HMM Cantonese Speech to Text System13 Feb 2024 0 repositories listed
-
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data8 Feb 2024 0 repositories listed
-
A Case Study on Filtering for End-to-End Speech Translation2 Feb 2024 0 repositories listed
-
Digits micro-model for accurate and secure transactions2 Feb 2024 0 repositories listed
-
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases1 Feb 2024 0 repositories listed
-
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks18 Jan 2024 0 repositories listed
-
OAVA: the open audio-visual archives aggregator16 Dec 2023 0 repositories listed
-
Revisiting the Entropy Semiring for Neural Speech Recognition13 Dec 2023 0 repositories listed
-
Efficient Monotonic Multihead Attention7 Dec 2023 0 repositories listed
-
End-to-End Speech-to-Text Translation: A Survey2 Dec 2023 0 repositories listed
-
Multi-teacher Distillation for Multilingual Spelling Correction20 Nov 2023 0 repositories listed
-
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning3 Nov 2023 0 repositories listed
-
Toward Joint Language Modeling for Speech Units and Text12 Oct 2023 0 repositories listed
-
Improving Stability in Simultaneous Speech Translation: A Revision-Controllable Decoding Approach6 Oct 2023 0 repositories listed
-
Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer5 Oct 2023 0 repositories listed
-
AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR30 Sep 2023 0 repositories listed
-
Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing27 Sep 2023 0 repositories listed
-
Developing automatic verbatim transcripts for international multilingual meetings: an end-to-end solution27 Sep 2023 0 repositories listed
-
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models22 Sep 2023 0 repositories listed
-
SpeechAlign: a Framework for Speech Translation Alignment Evaluation20 Sep 2023 0 repositories listed
-
CoLLD: Contrastive Layer-to-layer Distillation for Compressing Multilingual Pre-trained Speech Encoders14 Sep 2023 0 repositories listed
-
PhantomSound: Black-Box, Query-Efficient Audio Adversarial Attack via Split-Second Phoneme Injection13 Sep 2023 0 repositories listed
-
N-gram Boosting: Improving Contextual Biasing with Normalized N-gram Targets4 Aug 2023 0 repositories listed
-
Improving RNN-Transducers with Acoustic LookAhead11 Jul 2023 0 repositories listed
-
On decoder-only architecture for speech-to-text and large language model integration8 Jul 2023 0 repositories listed
-
Performance Comparison of Pre-trained Models for Speech-to-Text in Turkish: Whisper-Small and Wav2Vec2-XLS-R-300M6 Jul 2023 0 repositories listed
-
Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture5 Jul 2023 0 repositories listed
-
AudioPaLM: A Large Language Model That Can Speak and Listen22 Jun 2023 0 repositories listed
-
Recent Advances in Direct Speech-to-text Translation20 Jun 2023 0 repositories listed
-
Open Brain AI. Automatic Language Assessment11 Jun 2023 0 repositories listed
-
Speech-to-Text Adapter and Speech-to-Entity Retriever Augmented LLMs for Speech Understanding8 Jun 2023 0 repositories listed
-
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation1 Jun 2023 0 repositories listed
-
Strategies for improving low resource speech to text translation relying on pre-trained ASR models31 May 2023 0 repositories listed
-
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions30 May 2023 0 repositories listed
-
CIF-PT: Bridging Speech and Text Representations for Spoken Language Understanding via Continuous Integrate-and-Fire Pre-Training27 May 2023 0 repositories listed
-
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation25 May 2023 0 repositories listed
-
Improving Metrics for Speech Translation22 May 2023 0 repositories listed
-
Application-Agnostic Language Modeling for On-Device ASR16 May 2023 0 repositories listed
-
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks4 May 2023 0 repositories listed
-
Improving Autoregressive NLP Tasks via Modular Linearized Attention17 Apr 2023 0 repositories listed
-
Enhancing Speech-to-Speech Translation with Multiple TTS Targets10 Apr 2023 0 repositories listed
-
Natural Language Robot Programming: NLP integrated with autonomous robotic grasping6 Apr 2023 0 repositories listed
-
31 Mar 2023 0 repositories listed
-
wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts6 Mar 2023 0 repositories listed
-
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages2 Mar 2023 0 repositories listed
-
Improving Medical Speech-to-Text Accuracy with Vision-Language Pre-training Model27 Feb 2023 0 repositories listed
-
PATCorrect: Non-autoregressive Phoneme-augmented Transformer for ASR Error Correction10 Feb 2023 0 repositories listed
-
Characterizing Financial Market Coverage using Artificial Intelligence7 Feb 2023 0 repositories listed
-
Using External Off-Policy Speech-To-Text Mappings in Contextual End-To-End Automated Speech Recognition6 Jan 2023 0 repositories listed
-
Pushing the performances of ASR models on English and Spanish accents22 Dec 2022 0 repositories listed
-
M3ST: Mix at Three Levels for Speech Translation7 Dec 2022 0 repositories listed
-
Handling and extracting key entities from customer conversations using Speech recognition and Named Entity recognition28 Nov 2022 0 repositories listed
-
Multilingual Speech Emotion Recognition With Multi-Gating Mechanism and Neural Architecture Search31 Oct 2022 0 repositories listed
-
Phonemic Representation and Transcription for Speech to Text Applications for Under-resourced Indigenous African Languages: The Case of Kiswahili29 Oct 2022 0 repositories listed
-
Named Entity Detection and Injection for Direct Speech Translation21 Oct 2022 0 repositories listed
-
Improving Semi-supervised End-to-end Automatic Speech Recognition using CycleGAN and Inter-domain Losses20 Oct 2022 0 repositories listed
-
Simple and Effective Unsupervised Speech Translation18 Oct 2022 0 repositories listed
-
CTC Alignments Improve Autoregressive Translation11 Oct 2022 0 repositories listed
-
Speech-to-Text and Evaluation of Multiple Machine Translation Systems1 Sep 2022 0 repositories listed
-
Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks25 Aug 2022 0 repositories listed
-
Improving Hypernasality Estimation with Automatic Speech Recognition in Cleft Palate Speech10 Aug 2022 0 repositories listed
-
Extending RNN-T-based speech recognition systems with emotion and language classification28 Jul 2022 0 repositories listed
-
RSD-GAN: Regularized Sobolev Defense GAN Against Speech-to-Text Adversarial Attacks14 Jul 2022 0 repositories listed
-
Findings of the Third Workshop on Automatic Simultaneous Translation1 Jul 2022 0 repositories listed
-
Language Model Augmented Monotonic Attention for Simultaneous Translation1 Jul 2022 0 repositories listed
-
Swiss German Speech to Text system evaluation1 Jul 2022 0 repositories listed
-
System Description on Automatic Simultaneous Translation Workshop1 Jul 2022 0 repositories listed
-
Developing a Speech Recognition System for Recognizing Tonal Speech Signals Using a Convolutional Neural Network17 Jun 2022 0 repositories listed
-
A Semi-Automated Live Interlingual Communication Workflow Featuring Intralingual Respeaking: Evaluation and Benchmarking1 Jun 2022 0 repositories listed
-
The Nós Project: Opening routes for the Galician language in the field of language technologies1 Jun 2022 0 repositories listed
-
Towards Large Vocabulary Kazakh-Russian Sign Language Dataset: KRSL-OnlineSchool1 Jun 2022 0 repositories listed
-
Clinical Dialogue Transcription Error Correction using Seq2Seq Models26 May 2022 0 repositories listed
-
Semantic-preserved Communication System for Highly Efficient Speech Transmission25 May 2022 0 repositories listed
-
SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation17 May 2022 0 repositories listed
-
Hearing voices at the National Library -- a speech corpus and acoustic model for the Swedish language6 May 2022 0 repositories listed
-
Design of a novel Korean learning application for efficient pronunciation correction4 May 2022 0 repositories listed
-
Learning Adaptive Segmentation Policy for End-to-End Simultaneous Translation1 May 2022 0 repositories listed
-
NAIST Simultaneous Speech-to-Text Translation System for IWSLT 20221 May 2022 0 repositories listed
-
The AISP-SJTU Simultaneous Translation System for IWSLT 20221 May 2022 0 repositories listed
-
The HW-TSC’s Simultaneous Speech Translation System for IWSLT 2022 Evaluation1 May 2022 0 repositories listed
-
WaBERT: A Low-resource End-to-end Model for Spoken Language Understanding and Speech-to-BERT Alignment22 Apr 2022 0 repositories listed
-
Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation6 Apr 2022 0 repositories listed
-
A Study of Gender Impact in Self-supervised Models for Speech-to-Text Systems4 Apr 2022 0 repositories listed
-
Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents3 Apr 2022 0 repositories listed
-
The MIT Voice Name System28 Mar 2022 0 repositories listed
-
XTREME-S: Evaluating Cross-lingual Speech Representations21 Mar 2022 0 repositories listed
-
A combined approach to the analysis of speech conversations in a contact center domain12 Mar 2022 0 repositories listed
-
Attacks as Defenses: Designing Robust Audio CAPTCHAs Using Attacks on Automatic Speech Recognition Systems10 Mar 2022 0 repositories listed
-
Which French speech recognition system for assistant robots?4 Mar 2022 0 repositories listed
-
Punctuation restoration in Swedish through fine-tuned KB-BERT14 Feb 2022 0 repositories listed
-
Semantic-aware Speech to Text Transmission with Redundancy Removal7 Feb 2022 0 repositories listed
-
Optimization of a Real-Time Wavelet-Based Algorithm for Improving Speech Intelligibility5 Feb 2022 0 repositories listed
-
Cross-modal Contrastive Learning for Speech Translation17 Dec 2021 0 repositories listed
-
Training end-to-end speech-to-text models on mobile phones7 Dec 2021 0 repositories listed
-
An Experiment on Speech-to-Text Translation Systems for Manipuri to English on Low Resource Setting1 Dec 2021 0 repositories listed
-
Impact of Microphone position Measurement Error on Multi Channel Distant Speech Recognition & Intelligibility1 Dec 2021 0 repositories listed
-
Improve Sinhala Speech Recognition Through e2e LF-MMI Model1 Dec 2021 0 repositories listed
-
Comparison of SVD and factorized TDNN approaches for speech to text13 Oct 2021 0 repositories listed