{"url":"/method/griffin-lim-algorithm","slug":"griffin-lim-algorithm","name":"Griffin-Lim Algorithm","full_name":"Griffin-Lim Algorithm","full_name_withheld":false,"description_markdown":"The **Griffin-Lim Algorithm (GLA)** is a phase reconstruction method based on the redundancy of the short-time Fourier transform. It promotes the consistency of a spectrogram by iterating two projections, where a spectrogram is said to be consistent when its inter-bin dependency owing to the redundancy of STFT is retained.  GLA is based only on the consistency and does not take any prior knowledge about the target signal into account. \r\n\r\nThis algorithm expects to recover a complex-valued spectrogram, which is consistent and maintains the given amplitude $\\mathbf{A}$, by the following alternative projection procedure:\r\n\r\n$$ \\mathbf{X}^{[m+1]} = P\\_{\\mathcal{C}}\\left(P\\_{\\mathcal{A}}\\left(\\mathbf{X}^{[m]}\\right)\\right) $$\r\n\r\nwhere $\\mathbf{X}$ is a complex-valued spectrogram updated through the iteration, $P\\_{\\mathcal{S}}$ is the metric projection onto a set $\\mathcal{S}$, and $m$ is the iteration index. Here, $\\mathcal{C}$ is the set of consistent spectrograms, and $\\mathcal{A}$ is the set of spectrograms whose amplitude is the same as the given one. The metric projections onto these sets $\\mathcal{C}$ and $\\mathcal{A}$ are given by:\r\n\r\n$$ P\\_{\\mathcal{C}}(\\mathbf{X}) = \\mathcal{GG}^{†}\\mathbf{X} $$\r\n$$ P\\_{\\mathcal{A}}(\\mathbf{X}) = \\mathbf{A} \\odot \\mathbf{X} \\oslash |\\mathbf{X}| $$\r\n\r\n\r\nwhere $\\mathcal{G}$ represents STFT, $\\mathcal{G}^{†}$ is the pseudo inverse of STFT (iSTFT), $\\odot$ and $\\oslash$ are element-wise multiplication and division, respectively, and division by zero is replaced by zero. GLA is obtained as an algorithm for the following optimization problem:\r\n\r\n$$ \\min\\_{\\mathbf{X}} || \\mathbf{X} - P\\_{\\mathcal{C}}\\left(\\mathbf{X}\\right) ||^{2}\\_{\\text{Fro}} \\text{ s.t. } \\mathbf{X} \\in \\mathcal{A} $$\r\n\r\nwhere $ || · ||\\_{\\text{Fro}}$ is the Frobenius norm. This equation minimizes the energy of the inconsistent components under the constraint on amplitude which must be equal to the given one. Although GLA has been widely utilized because of its simplicity, GLA often involves many iterations until it converges to a certain spectrogram and results in low reconstruction quality. This is because the cost function only requires the consistency, and the characteristics of the target signal are not taken into account.","description_state":"present","introduced_year":1984,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Phase Reconstruction","url":"/methods/category/phase-reconstruction","pwc_aliases":[]}],"n_papers_tagged":75,"archive_num_papers":75,"papers_newest_first":[{"paper":"/paper/very-attentive-tacotron-robust-and-unbounded","title":"Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech","date":"2024-10-29","arxiv_id":"2410.22179","n_code_links":1,"syntology":null},{"paper":null,"title":"Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach","date":"2024-09-10","arxiv_id":"2409.13734","n_code_links":0,"syntology":null},{"paper":null,"title":"Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems","date":"2024-09-04","arxiv_id":"2409.02517","n_code_links":0,"syntology":null},{"paper":null,"title":"DDFAD: Dataset Distillation Framework for Audio Data","date":"2024-07-15","arxiv_id":"2407.10446","n_code_links":0,"syntology":null},{"paper":null,"title":"Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation","date":"2024-04-03","arxiv_id":"2404.02592","n_code_links":0,"syntology":null},{"paper":null,"title":"GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model","date":"2024-02-09","arxiv_id":"2402.15516","n_code_links":0,"syntology":null},{"paper":null,"title":"An overview of text-to-speech systems and media applications","date":"2023-10-22","arxiv_id":"2310.14301","n_code_links":0,"syntology":null},{"paper":null,"title":"Energy-Based Models For Speech Synthesis","date":"2023-10-19","arxiv_id":"2310.12765","n_code_links":0,"syntology":null},{"paper":null,"title":"A Flexible Online Framework for Projection-Based STFT Phase Retrieval","date":"2023-09-13","arxiv_id":"2309.07043","n_code_links":0,"syntology":null},{"paper":null,"title":"The DeepZen Speech Synthesis System for Blizzard Challenge 2023","date":"2023-08-30","arxiv_id":"2308.15945","n_code_links":0,"syntology":null},{"paper":"/paper/multilingual-text-to-speech-synthesis-for","title":"Multilingual Text-to-Speech Synthesis for Turkic Languages Using Transliteration","date":"2023-05-25","arxiv_id":"2305.15749","n_code_links":1,"syntology":null},{"paper":null,"title":"A Virtual Simulation-Pilot Agent for Training of Air Traffic Controllers","date":"2023-04-16","arxiv_id":"2304.07842","n_code_links":0,"syntology":null},{"paper":null,"title":"ArmanTTS single-speaker Persian dataset","date":"2023-04-07","arxiv_id":"2304.03585","n_code_links":0,"syntology":null},{"paper":null,"title":"Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language","date":"2022-12-16","arxiv_id":"2212.08321","n_code_links":0,"syntology":null},{"paper":null,"title":"Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features","date":"2022-11-01","arxiv_id":"2211.00342","n_code_links":0,"syntology":null},{"paper":null,"title":"Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation","date":"2022-10-31","arxiv_id":"2210.17264","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Developing State-of-the-Art TTS Synthesisers for 13 Indian Languages with Signal Processing aided Alignments","date":"2022-10-31","arxiv_id":"2210.17153","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficiently Trained Low-Resource Mongolian Text-to-Speech System Based On FullConv-TTS","date":"2022-10-24","arxiv_id":"2211.01948","n_code_links":0,"syntology":null},{"paper":"/paper/facial-landmark-predictions-with-applications","title":"Facial Landmark Predictions with Applications to Metaverse","date":"2022-09-29","arxiv_id":"2209.14698","n_code_links":1,"syntology":null},{"paper":null,"title":"Beyond Griffin-Lim: Improved Iterative Phase Retrieval for Speech","date":"2022-05-11","arxiv_id":"2205.05496","n_code_links":0,"syntology":null},{"paper":null,"title":"Self-supervised learning for robust voice cloning","date":"2022-04-07","arxiv_id":"2204.03421","n_code_links":0,"syntology":null},{"paper":null,"title":"Singing-Tacotron: Global duration control attention and dynamic filter for End-to-end singing voice synthesis","date":"2022-02-16","arxiv_id":"2202.07907","n_code_links":0,"syntology":null},{"paper":null,"title":"Zero-Shot Long-Form Voice Cloning with Dynamic Convolution Attention","date":"2022-01-25","arxiv_id":"2201.10375","n_code_links":0,"syntology":null},{"paper":null,"title":"Word-Level Style Control for Expressive, Non-attentive Speech Synthesis","date":"2021-11-19","arxiv_id":"2111.10173","n_code_links":0,"syntology":null},{"paper":null,"title":"High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency","date":"2021-11-17","arxiv_id":"2111.09052","n_code_links":0,"syntology":null},{"paper":null,"title":"On-device neural speech synthesis","date":"2021-09-17","arxiv_id":"2109.08710","n_code_links":0,"syntology":null},{"paper":null,"title":"Neural Sequence-to-Sequence Speech Synthesis Using a Hidden Semi-Markov Model Based Structured Attention Mechanism","date":"2021-08-31","arxiv_id":"2108.13985","n_code_links":0,"syntology":null},{"paper":"/paper/neural-hmms-are-all-you-need-for-high-quality","title":"Neural HMMs are all you need (for high-quality attention-free TTS)","date":"2021-08-30","arxiv_id":"2108.13320","n_code_links":2,"syntology":null},{"paper":"/paper/one-tts-alignment-to-rule-them-all","title":"One TTS Alignment To Rule Them All","date":"2021-08-23","arxiv_id":"2108.10447","n_code_links":3,"syntology":null},{"paper":null,"title":"Using Deep Learning Techniques and Inferential Speech Statistics for AI Synthesised Speech Recognition","date":"2021-07-23","arxiv_id":"2107.11412","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/speech-synthesis","name":"Speech Synthesis","papers":45},{"task":"/task/text-to-speech","name":"Text to Speech","papers":44},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":44},{"task":"/task/text-to-speech-synthesis","name":"Text-To-Speech Synthesis","papers":15},{"task":"/task/decoder","name":"Decoder","papers":10},{"task":"/task/sentence","name":"Sentence","papers":6},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":5},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":5},{"task":"/task/voice-cloning","name":"Voice Cloning","papers":5},{"task":"/task/audio-synthesis","name":"Audio Synthesis","papers":4},{"task":null,"name":"GPU","papers":4},{"task":"/task/voice-conversion","name":"Voice Conversion","papers":4},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":4},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":3},{"task":"/task/expressive-speech-synthesis","name":"Expressive Speech Synthesis","papers":3},{"task":null,"name":"Generative Adversarial Network","papers":3},{"task":"/task/speaker-verification","name":"Speaker Verification","papers":3},{"task":"/task/all","name":"All","papers":2},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":2},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":2}],"tasks_shown":20,"n_tasks":55,"usage_by_year":[{"year":"2017","papers":5},{"year":"2018","papers":7},{"year":"2019","papers":7},{"year":"2020","papers":19},{"year":"2021","papers":14},{"year":"2022","papers":10},{"year":"2023","papers":7},{"year":"2024","papers":6}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/griffin-lim-algorithm"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}