{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tunet-a-block-online-bandwidth-extension","title":"TUNet: A Block-online Bandwidth Extension Model based on Transformers and Self-supervised Pretraining","arxiv_id":"2110.13492","date":"2021-10-26","proceeding":null,"authors":["Viet-Anh Nguyen","Anh H. T. Nguyen","Andy W. H. Khong"],"abstract":"We introduce a block-online variant of the temporal feature-wise linear modulation (TFiLM) model to achieve bandwidth extension. The proposed architecture simplifies the UNet backbone of the TFiLM to reduce inference time and employs an efficient transformer at the bottleneck to alleviate performance degradation. We also utilize self-supervised pretraining and data augmentation to enhance the quality of bandwidth extended signals and reduce the sensitivity with respect to downsampling methods. Experiment results on the VCTK dataset show that the proposed method outperforms several recent baselines in both intrusive and non-intrusive metrics. Pretraining and filter augmentation also help stabilize and enhance the overall performance.","url_abs":"https://arxiv.org/abs/2110.13492v5","url_pdf":"https://arxiv.org/pdf/2110.13492v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tunet-a-block-online-bandwidth-extension","repo_url":"https://github.com/nxtproduct/tunet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"audio-super-resolution","task_name":"Audio Super-Resolution"},{"task_slug":"bandwidth-extension","task_name":"Bandwidth Extension"},{"task_slug":"sensitivity","task_name":"Sensitivity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-super-resolution-on-vctk-multi-speaker-1","task":"Audio Super-Resolution","dataset":"VCTK Multi-Speaker","model":"TUNet + MSM pre-training","rank_in_archive_order":3,"of":7,"metrics":{"Log-Spectral Distance":"1.28"},"uses_additional_data":false},{"leaderboard":"/sota/audio-super-resolution-on-vctk-multi-speaker-1","task":"Audio Super-Resolution","dataset":"VCTK Multi-Speaker","model":"TUNet","rank_in_archive_order":4,"of":7,"metrics":{"Log-Spectral Distance":"1.36"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2110.13492","atlas_url":"https://app.syntology.ai/?focus=2110.13492","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}