{"url":"/method/cbhg","slug":"cbhg","name":"CBHG","full_name":"CBHG","full_name_withheld":false,"description_markdown":"**CBHG** is a building block used in the [Tacotron](https://paperswithcode.com/method/tacotron) text-to-speech model. It consists of a bank of 1-D convolutional filters, followed by highway networks and a bidirectional gated recurrent unit ([BiGRU](https://paperswithcode.com/method/bigru)). \r\n\r\nThe module is used to extract representations from sequences. The input sequence is first\r\nconvolved with $K$ sets of 1-D convolutional filters, where the $k$-th set contains $C\\_{k}$ filters of width $k$ (i.e. $k = 1, 2, \\dots , K$). These filters explicitly model local and contextual information (akin to modeling unigrams, bigrams, up to K-grams). The [convolution](https://paperswithcode.com/method/convolution) outputs are stacked together and further max pooled along time to increase local invariances. A stride of 1 is used to  preserve the original time resolution. The processed sequence is further passed to a few fixed-width 1-D convolutions, whose outputs are added with the original input sequence via residual connections. [Batch normalization](https://paperswithcode.com/method/batch-normalization) is used for all convolutional layers. The convolution outputs are fed into a multi-layer [highway network](https://paperswithcode.com/method/highway-network) to extract high-level features. Finally, a bidirectional [GRU](https://paperswithcode.com/method/gru) RNN is stacked on top to extract sequential features from both forward and backward context.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Tacotron: Towards End-to-End Speech Synthesis","paper":"/paper/tacotron-towards-end-to-end-speech-synthesis","first_author":"Yuxuan Wang","n_authors":14,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/tacotron-towards-end-to-end-speech-synthesis"},"source":{"url":"http://arxiv.org/abs/1703.10135v2","title":"Tacotron: Towards End-to-End Speech Synthesis","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/mozilla/TTS/blob/3cbf9052f78ac40f025cdf598437eabc9bda9298/layers/tacotron.py#L93","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Skip Connection Blocks","url":"/methods/category/skip-connection-blocks","pwc_aliases":[]},{"area":"Sequential","area_id":"sequential","collection":"Sequential Blocks","url":"/methods/category/sequential-blocks","pwc_aliases":[]},{"area":"Audio","area_id":"audio","collection":"Speech Synthesis Blocks","url":"/methods/category/speech-synthesis-blocks","pwc_aliases":[]}],"n_papers_tagged":65,"archive_num_papers":65,"papers_newest_first":[{"paper":"/paper/very-attentive-tacotron-robust-and-unbounded","title":"Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech","date":"2024-10-29","arxiv_id":"2410.22179","n_code_links":1,"syntology":null},{"paper":null,"title":"Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach","date":"2024-09-10","arxiv_id":"2409.13734","n_code_links":0,"syntology":null},{"paper":null,"title":"Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems","date":"2024-09-04","arxiv_id":"2409.02517","n_code_links":0,"syntology":null},{"paper":null,"title":"Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation","date":"2024-04-03","arxiv_id":"2404.02592","n_code_links":0,"syntology":null},{"paper":null,"title":"An overview of text-to-speech systems and media applications","date":"2023-10-22","arxiv_id":"2310.14301","n_code_links":0,"syntology":null},{"paper":null,"title":"Energy-Based Models For Speech Synthesis","date":"2023-10-19","arxiv_id":"2310.12765","n_code_links":0,"syntology":null},{"paper":null,"title":"The DeepZen Speech Synthesis System for Blizzard Challenge 2023","date":"2023-08-30","arxiv_id":"2308.15945","n_code_links":0,"syntology":null},{"paper":"/paper/multilingual-text-to-speech-synthesis-for","title":"Multilingual Text-to-Speech Synthesis for Turkic Languages Using Transliteration","date":"2023-05-25","arxiv_id":"2305.15749","n_code_links":1,"syntology":null},{"paper":null,"title":"A Virtual Simulation-Pilot Agent for Training of Air Traffic Controllers","date":"2023-04-16","arxiv_id":"2304.07842","n_code_links":0,"syntology":null},{"paper":null,"title":"ArmanTTS single-speaker Persian dataset","date":"2023-04-07","arxiv_id":"2304.03585","n_code_links":0,"syntology":null},{"paper":null,"title":"Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language","date":"2022-12-16","arxiv_id":"2212.08321","n_code_links":0,"syntology":null},{"paper":null,"title":"Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features","date":"2022-11-01","arxiv_id":"2211.00342","n_code_links":0,"syntology":null},{"paper":null,"title":"Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation","date":"2022-10-31","arxiv_id":"2210.17264","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Developing State-of-the-Art TTS Synthesisers for 13 Indian Languages with Signal Processing aided Alignments","date":"2022-10-31","arxiv_id":"2210.17153","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficiently Trained Low-Resource Mongolian Text-to-Speech System Based On FullConv-TTS","date":"2022-10-24","arxiv_id":"2211.01948","n_code_links":0,"syntology":null},{"paper":"/paper/facial-landmark-predictions-with-applications","title":"Facial Landmark Predictions with Applications to Metaverse","date":"2022-09-29","arxiv_id":"2209.14698","n_code_links":1,"syntology":null},{"paper":null,"title":"Self-supervised learning for robust voice cloning","date":"2022-04-07","arxiv_id":"2204.03421","n_code_links":0,"syntology":null},{"paper":null,"title":"Singing-Tacotron: Global duration control attention and dynamic filter for End-to-end singing voice synthesis","date":"2022-02-16","arxiv_id":"2202.07907","n_code_links":0,"syntology":null},{"paper":null,"title":"Zero-Shot Long-Form Voice Cloning with Dynamic Convolution Attention","date":"2022-01-25","arxiv_id":"2201.10375","n_code_links":0,"syntology":null},{"paper":null,"title":"Word-Level Style Control for Expressive, Non-attentive Speech Synthesis","date":"2021-11-19","arxiv_id":"2111.10173","n_code_links":0,"syntology":null},{"paper":null,"title":"High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency","date":"2021-11-17","arxiv_id":"2111.09052","n_code_links":0,"syntology":null},{"paper":null,"title":"On-device neural speech synthesis","date":"2021-09-17","arxiv_id":"2109.08710","n_code_links":0,"syntology":null},{"paper":null,"title":"Neural Sequence-to-Sequence Speech Synthesis Using a Hidden Semi-Markov Model Based Structured Attention Mechanism","date":"2021-08-31","arxiv_id":"2108.13985","n_code_links":0,"syntology":null},{"paper":"/paper/neural-hmms-are-all-you-need-for-high-quality","title":"Neural HMMs are all you need (for high-quality attention-free TTS)","date":"2021-08-30","arxiv_id":"2108.13320","n_code_links":2,"syntology":null},{"paper":"/paper/one-tts-alignment-to-rule-them-all","title":"One TTS Alignment To Rule Them All","date":"2021-08-23","arxiv_id":"2108.10447","n_code_links":3,"syntology":null},{"paper":null,"title":"Using Deep Learning Techniques and Inferential Speech Statistics for AI Synthesised Speech Recognition","date":"2021-07-23","arxiv_id":"2107.11412","n_code_links":0,"syntology":null},{"paper":null,"title":"AI based Presentation Creator With Customized Audio Content Delivery","date":"2021-06-27","arxiv_id":"2106.14213","n_code_links":0,"syntology":null},{"paper":null,"title":"Ctrl-P: Temporal Control of Prosodic Variation for Speech Synthesis","date":"2021-06-15","arxiv_id":"2106.08352","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring emotional prototypes in a high dimensional TTS latent space","date":"2021-05-05","arxiv_id":"2105.01891","n_code_links":0,"syntology":null},{"paper":null,"title":"VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention","date":"2021-02-12","arxiv_id":"2102.06431","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/speech-synthesis","name":"Speech Synthesis","papers":43},{"task":"/task/text-to-speech","name":"Text to Speech","papers":41},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":41},{"task":"/task/text-to-speech-synthesis","name":"Text-To-Speech Synthesis","papers":15},{"task":"/task/decoder","name":"Decoder","papers":10},{"task":"/task/sentence","name":"Sentence","papers":6},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":5},{"task":"/task/voice-cloning","name":"Voice Cloning","papers":5},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":4},{"task":"/task/voice-conversion","name":"Voice Conversion","papers":4},{"task":"/task/audio-synthesis","name":"Audio Synthesis","papers":3},{"task":"/task/expressive-speech-synthesis","name":"Expressive Speech Synthesis","papers":3},{"task":null,"name":"GPU","papers":3},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":3},{"task":"/task/all","name":"All","papers":2},{"task":null,"name":"CPU","papers":2},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":2},{"task":"/task/diversity","name":"Diversity","papers":2},{"task":null,"name":"Generative Adversarial Network","papers":2},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":2}],"tasks_shown":20,"n_tasks":49,"usage_by_year":[{"year":"2017","papers":4},{"year":"2018","papers":6},{"year":"2019","papers":7},{"year":"2020","papers":16},{"year":"2021","papers":13},{"year":"2022","papers":9},{"year":"2023","papers":6},{"year":"2024","papers":4}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cbhg"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}