{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-film-capturing-long-range-sequence-1","title":"Temporal FiLM: Capturing Long-Range Sequence Dependencies with Feature-Wise Modulations.","arxiv_id":null,"date":"2019-12-01","proceeding":"NeurIPS 2019 12","authors":["Sawyer Birnbaum","Volodymyr Kuleshov","Zayd Enam","Pang Wei W. Koh","Stefano Ermon"],"abstract":"Learning representations that accurately capture long-range dependencies in sequential inputs --- including text, audio, and genomic data --- is a key problem in deep learning. Feed-forward convolutional models capture only feature interactions within finite receptive fields while recurrent architectures can be slow and difficult to train due to vanishing gradients. Here, we propose Temporal Feature-Wise Linear Modulation (TFiLM) --- a novel architectural component inspired by adaptive batch normalization and its extensions --- that uses a recurrent neural network to alter the activations of a convolutional model. This approach expands the receptive field of convolutional sequence models with minimal computational overhead. Empirically, we find that TFiLM significantly improves the learning speed and accuracy of feed-forward neural networks on a range of generative and discriminative learning tasks, including text classification and audio super-resolution.","url_abs":"http://papers.nips.cc/paper/9217-temporal-film-capturing-long-range-sequence-dependencies-with-feature-wise-modulations","url_pdf":"http://papers.nips.cc/paper/9217-temporal-film-capturing-long-range-sequence-dependencies-with-feature-wise-modulations.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporal-film-capturing-long-range-sequence-1","repo_url":"https://github.com/kuleshov/audio-super-res","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"audio-super-resolution","task_name":"Audio Super-Resolution"},{"task_slug":"super-resolution","task_name":"Super-Resolution"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-super-resolution-on-vctk-multi-speaker-1","task":"Audio Super-Resolution","dataset":"VCTK Multi-Speaker","model":"U-Net + TFiLM","rank_in_archive_order":6,"of":7,"metrics":{"Log-Spectral Distance":"1.8"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}