Methods › General › Loss Functions › CTC Loss
Connectionist Temporal Classification Loss
CTC Loss
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
A Connectionist Temporal Classification Loss, or CTC Loss, is designed for tasks where we need alignment between sequences, but where that alignment is difficult - e.g. aligning each character to its location in an audio file. It calculates a loss between a continuous (unsegmented) time series and a target sequence. It does this by summing over the probability of possible alignments of input to target, producing a loss value which is differentiable with respect to each input node. The alignment of input to target is assumed to be “many-to-one”, which limits the length of the target sequence such that it must be ≤ the input length.
Papers archive 2025-07-28
30 shown of 47, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation 2 Jun 2025 · 0 repositories · arXiv:2506.01503
-
Integrating Canonical Neural Units and Multi-Scale Training for Handwritten Text Recognition 24 Oct 2024 · 0 repositories · arXiv:2410.18374
-
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition 1 Sep 2024 · 0 repositories · arXiv:2409.00815
-
LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition 11 Aug 2024 · 1 repository · arXiv:2408.05769
-
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment 25 Jun 2024 · 0 repositories · arXiv:2406.17957
-
Best Practices for a Handwritten Text Recognition System 17 Apr 2024 · 1 repository · arXiv:2404.11339
-
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition 18 Mar 2024 · 0 repositories · arXiv:2403.11578
-
Key Frame Mechanism For Efficient Conformer Based End-to-end Speech Recognition 23 Oct 2023 · 1 repository · arXiv:2310.14954
-
Self-distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach 17 Aug 2023 · 1 repository · arXiv:2308.08806
-
Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition 12 Aug 2023 · 0 repositories · arXiv:2308.06547
-
Align With Purpose: Optimize Desired Properties in CTC Models with a General Plug-and-Play Framework 4 Jul 2023 · 0 repositories · arXiv:2307.01715
-
Improving Non-autoregressive Translation Quality with Pretrained Language Model, Embedding Distillation and Upsampling Strategy for CTC 10 Jun 2023 · 0 repositories · arXiv:2306.06345
-
INTapt: Information-Theoretic Adversarial Prompt Tuning for Enhanced Non-Native Speech Recognition 25 May 2023 · 0 repositories · arXiv:2305.16371
-
SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition 2 Dec 2022 · 1 repository · arXiv:2212.01039Syntology ran 6 of 6 samples · 0 unverified
-
Weakly-supervised Fingerspelling Recognition in British Sign Language Videos 16 Nov 2022 · 1 repository · arXiv:2211.08954
-
Peak-First CTC: Reducing the Peak Latency of CTC Models by Applying Peak-First Regularization 7 Nov 2022 · 0 repositories · arXiv:2211.03284
-
Temporal superimposed crossover module for effective continuous sign language 7 Nov 2022 · 1 repository · arXiv:2211.03387
-
TrimTail: Low-Latency Streaming ASR with Simple but Effective Spectrogram-Level Length Penalty 1 Nov 2022 · 1 repository · arXiv:2211.00522Syntology ran 3 of 5 samples · 2 unverified
-
Uconv-Conformer: High Reduction of Input Sequence Length for End-to-End Speech Recognition 16 Aug 2022 · 0 repositories · arXiv:2208.07657
-
Leveraging Acoustic Contextual Representation by Audio-textual Cross-modal Learning for Conversational ASR 3 Jul 2022 · 0 repositories · arXiv:2207.01039
-
Multi-scale temporal network for continuous sign language recognition 8 Apr 2022 · 0 repositories · arXiv:2204.03864
-
Non-Autoregressive ASR with Self-Conditioned Folded Encoders 17 Feb 2022 · 0 repositories · arXiv:2202.08474
-
PM-MMUT: Boosted Phone-Mask Data Augmentation using Multi-Modeling Unit Training for Phonetic-Reduction-Robust E2E Speech Recognition 13 Dec 2021 · 0 repositories · arXiv:2112.06721
-
Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates 27 Sep 2021 · 1 repository · arXiv:2109.12804
-
Golos: Russian Dataset for Speech Research 18 Jun 2021 · 2 repositories · arXiv:2106.10161Syntology ran 4 of 5 samples · 1 unverified · 3 pointer-only (licence)
-
Multi-Speaker ASR Combining Non-Autoregressive Conformer CTC and Conditional Speaker Chain 16 Jun 2021 · 1 repository · arXiv:2106.08595
-
Why does CTC result in peaky behavior? 31 May 2021 · 1 repository · arXiv:2105.14849
-
Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions 6 Apr 2021 · 0 repositories · arXiv:2104.02724
-
Multiple-hypothesis CTC-based semi-supervised adaptation of end-to-end speech recognition 29 Mar 2021 · 0 repositories · arXiv:2103.15515
-
Intermediate Loss Regularization for CTC-based Speech Recognition 5 Feb 2021 · 0 repositories · arXiv:2102.03216
Tasks archive 2025-07-28
20 shown of 42 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections