{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/frequency-domain-transformer-networks-for","title":"Frequency Domain Transformer Networks for Video Prediction","arxiv_id":"1903.00271","date":"2019-03-01","proceeding":null,"authors":["Hafez Farazi","Sven Behnke"],"abstract":"The task of video prediction is forecasting the next frames given some\nprevious frames. Despite much recent progress, this task is still challenging\nmainly due to high nonlinearity in the spatial domain. To address this issue,\nwe propose a novel architecture, Frequency Domain Transformer Network (FDTN),\nwhich is an end-to-end learnable model that estimates and uses the\ntransformations of the signal in the frequency domain. Experimental evaluations\nshow that this approach can outperform some widely used video prediction\nmethods like Video Ladder Network (VLN) and Predictive Gated Pyramids (PGP).","url_abs":"http://arxiv.org/abs/1903.00271v1","url_pdf":"http://arxiv.org/pdf/1903.00271v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"frequency-domain-transformer-networks-for","repo_url":"https://github.com/AIS-Bonn/FreqNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}