Browse State-of-the-Art › Speech-to-Text › Papers, page 2
Speech-to-Text
Papers archive 2025-07-28
archive papers tagged: 403 · with a code link: 129 · where Syntology ran a sample: 24 (18 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (24 of 403 tagged: 18 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument)
Page 2 of 5: papers 101 to 200 of 403, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
7 Sep 2021 1 repository listed
-
6 Sep 2021 1 repository listed
-
1 Aug 2021 1 repository listed
-
6 Jul 2021 1 repository listed
-
24 Jun 2021 1 repository listed
-
11 May 2021 1 repository listed
-
21 Apr 2021 1 repository listed
-
5 Apr 2021 1 repository listed
-
3 Dec 2020 1 repository listed
-
1 Dec 2020 1 repository listed
-
25 Nov 2020 1 repository listed
-
1 Nov 2020 1 repository listed
-
21 Oct 2020 1 repository listed
-
22 Sep 2020 1 repository listed
-
21 Sep 2020 1 repository listed
-
21 Sep 2020 1 repository listed
-
5 Aug 2020 1 repository listed
-
4 Feb 2020 1 repository listed Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
18 Jan 2020 1 repository listed Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 11 harvested samples)
-
1 Jan 2020 1 repository listed Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
16 Dec 2019 1 repository listed
-
6 Dec 2019 1 repository listed
-
29 Nov 2019 1 repository listed
-
12 Apr 2019 1 repository listed
-
5 Sep 2018 1 repository listed
-
12 Feb 2018 1 repository listed
-
9 Feb 2018 1 repository listed
-
6 Dec 2016 1 repository listed Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples)
-
20 Sep 2016 1 repository listed
-
An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments14 Jul 2025 0 repositories listed
-
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization20 Jun 2025 0 repositories listed
-
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data19 Jun 2025 0 repositories listed
-
I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs17 Jun 2025 0 repositories listed
-
S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Streaming Speech Generation11 Jun 2025 0 repositories listed
-
Advancing STT for Low-Resource Real-World Speech10 Jun 2025 0 repositories listed
-
Improving Language and Modality Transfer in Translation by Character-level Modeling30 May 2025 0 repositories listed
-
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios30 May 2025 0 repositories listed
-
Conversational Recommendation System using NLP and Sentiment Analysis17 May 2025 0 repositories listed
-
Acquisition of high-quality images for camera calibration in robotics applications via speech prompts15 Apr 2025 0 repositories listed
-
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect3 Apr 2025 0 repositories listed
-
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit26 Mar 2025 0 repositories listed
-
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation18 Mar 2025 0 repositories listed
-
Focusing Robot Open-Ended Reinforcement Learning Through Users' Purposes16 Mar 2025 0 repositories listed
-
Telephone Surveys Meet Conversational AI: Evaluating a LLM-Based Telephone Survey System at Scale27 Feb 2025 0 repositories listed
-
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision26 Feb 2025 0 repositories listed
-
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM24 Feb 2025 0 repositories listed
-
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation24 Feb 2025 0 repositories listed
-
Speech to Speech Translation with Translatotron: A State of the Art Review9 Feb 2025 0 repositories listed
-
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation1 Feb 2025 0 repositories listed
-
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction10 Jan 2025 0 repositories listed
-
Existential Crisis: A Social Robot's Reason for Being6 Jan 2025 0 repositories listed
-
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison4 Jan 2025 0 repositories listed
-
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages31 Dec 2024 0 repositories listed
-
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?24 Dec 2024 0 repositories listed
-
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation11 Dec 2024 0 repositories listed
-
Representation Purification for End-to-End Speech Translation5 Dec 2024 0 repositories listed
-
Leveraging Virtual Reality and AI Tutoring for Language Learning: A Case Study of a Virtual Campus Environment with OpenAI GPT Integration with Unity 3D19 Nov 2024 0 repositories listed
-
Whisper Finetuning on Nepali Language19 Nov 2024 0 repositories listed
-
Isochrony-Controlled Speech-to-Text Translation: A study on translating from Sino-Tibetan to Indo-European Languages11 Nov 2024 0 repositories listed
-
NeKo: Toward Post Recognition Generative Correction Large Language Models with Task-Oriented Experts8 Nov 2024 0 repositories listed
-
CUIfy the XR: An Open-Source Package to Embed LLM-powered Conversational Agents in XR7 Nov 2024 0 repositories listed
-
LASER: Attention with Exponential Transformation5 Nov 2024 0 repositories listed
-
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?31 Oct 2024 0 repositories listed
-
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization29 Oct 2024 0 repositories listed
-
A Survey on Speech Large Language Models24 Oct 2024 0 repositories listed
-
Contextual Biasing to Improve Domain-specific Custom Vocabulary Audio Transcription without Explicit Fine-Tuning of Whisper Model24 Oct 2024 0 repositories listed
-
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum18 Oct 2024 0 repositories listed
-
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck15 Oct 2024 0 repositories listed
-
Unsupervised Data Validation Methods for Efficient Model Training10 Oct 2024 0 repositories listed
-
Transducer Consistency Regularization for Speech to Text Applications9 Oct 2024 0 repositories listed
-
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems3 Oct 2024 0 repositories listed
-
Unveiling the Role of Pretraining in Direct Speech Translation26 Sep 2024 0 repositories listed
-
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not25 Sep 2024 0 repositories listed
-
On the Feasibility of Fully AI-automated Vishing Attacks20 Sep 2024 0 repositories listed
-
Toward Automated Clinical Transcriptions20 Sep 2024 0 repositories listed
-
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text17 Sep 2024 0 repositories listed
-
Evaluation of real-time transcriptions using end-to-end ASR models9 Sep 2024 0 repositories listed
-
5 Sep 2024 0 repositories listed
-
AI-Based IVR20 Aug 2024 0 repositories listed
-
CMU's IWSLT 2024 Simultaneous Speech Translation System14 Aug 2024 0 repositories listed
-
AI-Powered Immersive Assistance for Interactive Task Execution in Industrial Environments12 Jul 2024 0 repositories listed
-
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks10 Jul 2024 0 repositories listed
-
Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation4 Jul 2024 0 repositories listed
-
Investigating Decoder-only Large Language Models for Speech-to-text Translation3 Jul 2024 0 repositories listed
-
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders2 Jul 2024 0 repositories listed
-
NAIST Simultaneous Speech Translation System for IWSLT 202430 Jun 2024 0 repositories listed
-
Transferable speech-to-text large language model alignment module19 Jun 2024 0 repositories listed
-
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving16 Jun 2024 0 repositories listed
-
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models13 Jun 2024 0 repositories listed
-
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?11 Jun 2024 0 repositories listed
-
Synthetic Query Generation using Large Language Models for Virtual Assistants10 Jun 2024 0 repositories listed
-
VR-GPT: Visual Language Model for Intelligent Virtual Reality Applications19 May 2024 0 repositories listed
-
Semantic MIMO Systems for Speech-to-Text Transmission13 May 2024 0 repositories listed
-
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)2 May 2024 0 repositories listed
-
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair18 Apr 2024 0 repositories listed
-
NaturalTurn: A Method to Segment Transcripts into Naturalistic Conversational Turns22 Mar 2024 0 repositories listed
-
Rich Semantic Knowledge Enhanced Large Language Models for Few-shot Chinese Spell Checking13 Mar 2024 0 repositories listed
-
Robust Semantic Communications for Speech Transmission8 Mar 2024 0 repositories listed
-
Compact Speech Translation Models via Discrete Speech Units Pretraining29 Feb 2024 0 repositories listed
-
Direct Punjabi to English speech translation using discrete units25 Feb 2024 0 repositories listed
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.