{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/inr-v-a-continuous-representation-space-for","title":"INR-V: A Continuous Representation Space for Video-based Generative Tasks","arxiv_id":"2210.16579","date":"2022-10-29","proceeding":null,"authors":["Bipasha Sen","Aditya Agarwal","Vinay P Namboodiri","C. V. Jawahar"],"abstract":"Generating videos is a complex task that is accomplished by generating a set of temporally coherent images frame-by-frame. This limits the expressivity of videos to only image-based operations on the individual video frames needing network designs to obtain temporally coherent trajectories in the underlying image space. We propose INR-V, a video representation network that learns a continuous space for video-based generative tasks. INR-V parameterizes videos using implicit neural representations (INRs), a multi-layered perceptron that predicts an RGB value for each input pixel location of the video. The INR is predicted using a meta-network which is a hypernetwork trained on neural representations of multiple video instances. Later, the meta-network can be sampled to generate diverse novel videos enabling many downstream video-based generative tasks. Interestingly, we find that conditional regularization and progressive weight initialization play a crucial role in obtaining INR-V. The representation space learned by INR-V is more expressive than an image space showcasing many interesting properties not possible with the existing works. For instance, INR-V can smoothly interpolate intermediate videos between known video instances (such as intermediate identities, expressions, and poses in face videos). It can also in-paint missing portions in videos to recover temporally coherent full videos. In this work, we evaluate the space learned by INR-V on diverse generative tasks such as video interpolation, novel video generation, video inversion, and video inpainting against the existing baselines. INR-V significantly outperforms the baselines on several of these demonstrated tasks, clearly showcasing the potential of the proposed representation space.","url_abs":"https://arxiv.org/abs/2210.16579v2","url_pdf":"https://arxiv.org/pdf/2210.16579v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"inr-v-a-continuous-representation-space-for","repo_url":"https://github.com/bipashasen/INRV","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"inr-v-a-continuous-representation-space-for","repo_url":"https://github.com/MindSpore-scientific-2/code-5/tree/main/INR-Implicit-Neural-Representations-with-Periodic-Activation-Functions","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"inr-v-a-continuous-representation-space-for","repo_url":"https://github.com/MindSpore-scientific/code-11/tree/main/INR-Implicit-Neural-Representations-with-Periodic-Activation-Functions","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"video-generation","task_name":"Video Generation"},{"task_slug":"video-inpainting","task_name":"Video Inpainting"}],"methods":[{"method_slug":"hypernetwork","method_name":"HyperNetwork"},{"method_slug":"pixel-prediction","method_name":"Inpainting"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-generation-on-how2sign","task":"Video Generation","dataset":"How2Sign","model":"INR-V","rank_in_archive_order":1,"of":1,"metrics":{"FVD16":"144"},"uses_additional_data":false},{"leaderboard":"/sota/video-inpainting-on-how2sign","task":"Video Inpainting","dataset":"How2Sign","model":"INR-V","rank_in_archive_order":1,"of":1,"metrics":{"L1 error":"4.51"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.16579","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}