{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/real-time-video-super-resolution-with-spatio","title":"Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation","arxiv_id":"1611.05250","date":"2016-11-16","proceeding":"CVPR 2017 7","authors":["Jose Caballero","Christian Ledig","Andrew Aitken","Alejandro Acosta","Johannes Totz","Zehan Wang","Wenzhe Shi"],"abstract":"Convolutional neural networks have enabled accurate image super-resolution in\nreal-time. However, recent attempts to benefit from temporal correlations in\nvideo super-resolution have been limited to naive or inefficient architectures.\nIn this paper, we introduce spatio-temporal sub-pixel convolution networks that\neffectively exploit temporal redundancies and improve reconstruction accuracy\nwhile maintaining real-time speed. Specifically, we discuss the use of early\nfusion, slow fusion and 3D convolutions for the joint processing of multiple\nconsecutive video frames. We also propose a novel joint motion compensation and\nvideo super-resolution algorithm that is orders of magnitude more efficient\nthan competing methods, relying on a fast multi-resolution spatial transformer\nmodule that is end-to-end trainable. These contributions provide both higher\naccuracy and temporally more consistent videos, which we confirm qualitatively\nand quantitatively. Relative to single-frame models, spatio-temporal networks\ncan either reduce the computational cost by 30% whilst maintaining the same\nquality or provide a 0.2dB gain for a similar computational cost. Results on\npublicly available datasets demonstrate that the proposed algorithms surpass\ncurrent state-of-the-art performance in both accuracy and efficiency.","url_abs":"http://arxiv.org/abs/1611.05250v2","url_pdf":"http://arxiv.org/pdf/1611.05250v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"motion-compensation","task_name":"Motion Compensation"},{"task_slug":"video-super-resolution","task_name":"Video Super-Resolution"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-super-resolution-on-msu-video-upscalers","task":"Video Super-Resolution","dataset":"MSU Video Upscalers: Quality Enhancement","model":"VESPCN","rank_in_archive_order":39,"of":48,"metrics":{"PSNR":"26.92","SSIM":"0.932","VMAF":"53.96"},"uses_additional_data":false},{"leaderboard":"/sota/video-super-resolution-on-vid4-4x-upscaling","task":"Video Super-Resolution","dataset":"Vid4 - 4x upscaling","model":"VESPCN","rank_in_archive_order":19,"of":27,"metrics":{"MOVIE":"5.82","PSNR":"25.35","SSIM":"0.7557"},"uses_additional_data":false},{"leaderboard":"/sota/video-super-resolution-on-vid4-4x-upscaling","task":"Video Super-Resolution","dataset":"Vid4 - 4x upscaling","model":"bicubic","rank_in_archive_order":24,"of":27,"metrics":{"MOVIE":"9.31","PSNR":"23.82","SSIM":"0.6548"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1611.05250","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}