{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-film-capturing-long-range-sequence","title":"Temporal FiLM: Capturing Long-Range Sequence Dependencies with Feature-Wise Modulations","arxiv_id":"1909.06628","date":"2019-09-14","proceeding":null,"authors":["Sawyer Birnbaum","Volodymyr Kuleshov","Zayd Enam","Pang Wei Koh","Stefano Ermon"],"abstract":"Learning representations that accurately capture long-range dependencies in sequential inputs -- including text, audio, and genomic data -- is a key problem in deep learning. Feed-forward convolutional models capture only feature interactions within finite receptive fields while recurrent architectures can be slow and difficult to train due to vanishing gradients. Here, we propose Temporal Feature-Wise Linear Modulation (TFiLM) -- a novel architectural component inspired by adaptive batch normalization and its extensions -- that uses a recurrent neural network to alter the activations of a convolutional model. This approach expands the receptive field of convolutional sequence models with minimal computational overhead. Empirically, we find that TFiLM significantly improves the learning speed and accuracy of feed-forward neural networks on a range of generative and discriminative learning tasks, including text classification and audio super-resolution","url_abs":"https://arxiv.org/abs/1909.06628v3","url_pdf":"https://arxiv.org/pdf/1909.06628v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporal-film-capturing-long-range-sequence","repo_url":"https://github.com/leolya/Audio-Super-Resolution-Tensorflow2.0-TFiLM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"audio-super-resolution","task_name":"Audio Super-Resolution"},{"task_slug":"super-resolution","task_name":"Super-Resolution"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-super-resolution-on-piano-1","task":"Audio Super-Resolution","dataset":"Piano","model":"U-Net + TFiLM","rank_in_archive_order":2,"of":3,"metrics":{"Log-Spectral Distance":"2"},"uses_additional_data":false},{"leaderboard":"/sota/audio-super-resolution-on-voice-bank-corpus-1","task":"Audio Super-Resolution","dataset":"Voice Bank corpus (VCTK)","model":"U-Net + TFiLM","rank_in_archive_order":2,"of":3,"metrics":{"Log-Spectral Distance":"2.5"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1909.06628","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}