{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tdan-temporally-deformable-alignment-network","title":"TDAN: Temporally Deformable Alignment Network for Video Super-Resolution","arxiv_id":"1812.02898","date":"2018-12-07","proceeding":null,"authors":["Yapeng Tian","Yulun Zhang","Yun Fu","Chenliang Xu"],"abstract":"Video super-resolution (VSR) aims to restore a photo-realistic\nhigh-resolution (HR) video frame from both its corresponding low-resolution\n(LR) frame (reference frame) and multiple neighboring frames (supporting\nframes). Due to varying motion of cameras or objects, the reference frame and\neach support frame are not aligned. Therefore, temporal alignment is a\nchallenging yet important problem for VSR. Previous VSR methods usually utilize\noptical flow between the reference frame and each supporting frame to wrap the\nsupporting frame for temporal alignment. Therefore, the performance of these\nimage-level wrapping-based models will highly depend on the prediction accuracy\nof optical flow, and inaccurate optical flow will lead to artifacts in the\nwrapped supporting frames, which also will be propagated into the reconstructed\nHR video frame. To overcome the limitation, in this paper, we propose a\ntemporal deformable alignment network (TDAN) to adaptively align the reference\nframe and each supporting frame at the feature level without computing optical\nflow. The TDAN uses features from both the reference frame and each supporting\nframe to dynamically predict offsets of sampling convolution kernels. By using\nthe corresponding kernels, TDAN transforms supporting frames to align with the\nreference frame. To predict the HR video frame, a reconstruction network taking\naligned frames and the reference frame is utilized. Experimental results\ndemonstrate the effectiveness of the proposed TDAN-based VSR model.","url_abs":"http://arxiv.org/abs/1812.02898v1","url_pdf":"http://arxiv.org/pdf/1812.02898v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tdan-temporally-deformable-alignment-network","repo_url":"https://github.com/YapengTian/TDAN-VSR-CVPR-2020","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"tdan-temporally-deformable-alignment-network","repo_url":"https://github.com/YapengTian/TDAN_VSR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"super-resolution","task_name":"Super-Resolution"},{"task_slug":"video-super-resolution","task_name":"Video Super-Resolution"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-super-resolution-on-msu-vsr-benchmark","task":"Video Super-Resolution","dataset":"MSU Video Super Resolution Benchmark: Detail Restoration","model":"TDAN","rank_in_archive_order":12,"of":32,"metrics":{"1 - LPIPS":"0.721","ERQAv1.0":"0.706","FPS":"0.493","PSNR":"30.244","QRCRv1.0":"0.609","SSIM":"0.883","Subjective score":"5.454"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.02898","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}