{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-attention-for-audio-super-resolution","title":"Self-Attention for Audio Super-Resolution","arxiv_id":"2108.11637","date":"2021-08-26","proceeding":null,"authors":["Nathanaël Carraz Rakotonirina"],"abstract":"Convolutions operate only locally, thus failing to model global interactions. Self-attention is, however, able to learn representations that capture long-range dependencies in sequences. We propose a network architecture for audio super-resolution that combines convolution and self-attention. Attention-based Feature-Wise Linear Modulation (AFiLM) uses self-attention mechanism instead of recurrent neural networks to modulate the activations of the convolutional model. Extensive experiments show that our model outperforms existing approaches on standard benchmarks. Moreover, it allows for more parallelization resulting in significantly faster training.","url_abs":"https://arxiv.org/abs/2108.11637v1","url_pdf":"https://arxiv.org/pdf/2108.11637v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-attention-for-audio-super-resolution","repo_url":"https://github.com/ncarraz/AFILM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"audio-super-resolution","task_name":"Audio Super-Resolution"},{"task_slug":"super-resolution","task_name":"Super-Resolution"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-super-resolution-on-piano-1","task":"Audio Super-Resolution","dataset":"Piano","model":"U-Net + AFiLM","rank_in_archive_order":1,"of":3,"metrics":{"Log-Spectral Distance":"1.5"},"uses_additional_data":false},{"leaderboard":"/sota/audio-super-resolution-on-vctk-multi-speaker-1","task":"Audio Super-Resolution","dataset":"VCTK Multi-Speaker","model":"U-Net + AFiLM","rank_in_archive_order":5,"of":7,"metrics":{"Log-Spectral Distance":"1.7"},"uses_additional_data":false},{"leaderboard":"/sota/audio-super-resolution-on-voice-bank-corpus-1","task":"Audio Super-Resolution","dataset":"Voice Bank corpus (VCTK)","model":"U-Net + AFiLM","rank_in_archive_order":1,"of":3,"metrics":{"Log-Spectral Distance":"2.3"},"uses_additional_data":true}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}