Browse State-of-the-Art › Image to Video Generation
Image to Video Generation
38 papers with code · 0 benchmarks · 7 datasets archive 2025-07-28
Image to Video Generation refers to the task of generating a sequence of video frames based on a single still image or a set of still images. The goal is to produce a video that is coherent and consistent in terms of appearance, motion, and style, while also being temporally consistent, meaning that the generated video should look like a coherent sequence of frames that are temporally ordered. This task is typically tackled using deep generative models, such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs), that are trained on large datasets of videos. The models learn to generate plausible video frames that are conditioned on the input image, as well as on any other auxiliary information, such as a sound or text track.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 38 papers with code (85 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Jul 2022 4 repositories listedDrawing images of characters with desired poses is an essential but laborious task in anime production.
-
25 Nov 2023 3 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedWe then explore the impact of finetuning our base model on high-quality data and train a text-to-video model that is competitive with closed-source video generation.
-
29 Dec 2024 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedTo facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content.
-
5 Dec 2022 2 repositories listed3D reconstruction from a single 2D image was extensively covered in the literature but relies on depth supervision at training time, which limits its applicability.
-
21 Mar 2020 2 repositories listedMost existing aging methods are limited to changing the texture, overlooking transformations in head shape that occur during the human aging and growth process.
-
11 Mar 2019 2 repositories listedThis paper proposes the novel task of video generation conditioned on a SINGLE semantic label map, which provides a good balance between flexibility and quality in the generation process.
-
31 Mar 2025 1 repository listedOur approach introduces synthetic proxy images through two key innovations: (1) Dual-path score distillation: We employ a dual-path architecture to distill motion priors from both real and synthetic data, preserving…
-
14 Mar 2025 1 repository listedWe present Step-Video-TI2V, a state-of-the-art text-driven image-to-video generation model with 30B parameters, capable of generating videos up to 102 frames based on both text and image inputs.
-
5 Mar 2025 1 repository listedIn this work, we present DualDiff, a dual-branch conditional diffusion model designed to enhance driving scene generation across multiple views and video sequences.
-
2 Mar 2025 1 repository listed Syntology ran 6 of 15 samples · 9 unverified(3) With the above two-stage models excelling in motion controllability and degree, we decouple the relevant parameters associated with each type of motion ability and inject them into the base I2V-DM.
-
17 Feb 2025 1 repository listedAccurate and flexible world models are crucial for autonomous systems to understand their environment and predict future events.
-
11 Feb 2025 1 repository listedThe key idea is simple: factorize the text-to-video generation task into two separate easier tasks for diffusion step distillation, namely text-to-image generation and image-to-video generation.
-
14 Jan 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedWe highlight the effectiveness and efficiency of FramePainter across various of editing signals: it domainantly outperforms previous state-of-the-art methods with far less training data, achieving highly seamless and…
-
30 Dec 2024 1 repository listed Syntology ran 0 of 9 samples · 9 unverifiedTo address this, our VAE decoder is tasked with both latent-to-pixel conversion and the final denoising step, producing the clean result directly in pixel space.
-
13 Dec 2024 1 repository listedIn this paper, we present a comprehensive system, AniSora, designed for animation video generation, which includes a data processing pipeline, a controllable generation model, and an evaluation benchmark.
-
26 Nov 2024 1 repository listedWe propose a hierarchical training strategy to leverage frequency information for identity preservation, transforming a vanilla pre-trained video generation model into an IPT2V model.
-
20 Nov 2024 1 repository listedVideo generation has witnessed significant advancements, yet evaluating these models remains a challenge.
-
15 Nov 2024 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedRecent work on human animation usually involves audio, pose, or movement maps conditions, thereby achieves vivid animation quality.
-
24 Oct 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Current image watermarking methods are vulnerable to advanced image editing techniques enabled by large-scale text-to-image models.
-
4 Oct 2024 1 repository listed Syntology ran 7 of 10 samples · 3 unverifiedDiffusion models have revolutionized image generation, and their extension to video generation has shown promise.
-
27 Sep 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We present PhysGen, a novel image-to-video generation method that converts a single image and an input condition (e.
-
27 Aug 2024 1 repository listedIn this paper, we aim to investigate whether such priors derived from a generative process are suitable for video recognition, and eventually joint optimization of generation and recognition.
-
30 Jul 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To fulfill this gap, we present MMTrail, a large-scale multi-modality video-language dataset incorporating more than 20M trailer clips with visual captions, and 2M high-quality clips with multimodal captions.
-
22 Jun 2024 1 repository listedAlthough image composition based on diffusion models has been highly successful, it is not straightforward to extend the achievement to video object composition tasks, which not only exhibit corresponding interaction…
-
12 Jun 2024 1 repository listed Syntology ran 5 of 5 samples · 0 unverifiedVideo generation has many unique challenges beyond those of image generation.
-
25 Apr 2024 1 repository listedTo guide video generation with the additional image input, we propose a "repeat-and-slide" strategy that modulates the reverse denoising process, allowing the frozen diffusion model to synthesize a video frame-by-frame…
-
21 Mar 2024 1 repository listed Syntology ran 2 of 7 samples · 5 unverifiedAnyV2V can leverage any existing image editing tools to support an extensive array of video editing tasks, including prompt-based editing, reference-based style transfer, subject-driven editing, and identity…
-
21 Mar 2024 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedIn this study, we introduce a methodology for human image animation by leveraging a 3D human parametric model within a latent diffusion framework to enhance shape alignment and motion guidance in curernt human…
-
20 Mar 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)Existing open-source methods struggle to achieve comparable performance, often hindered by ineffective agent collaboration and inadequate training data quality.
-
13 Mar 2024 1 repository listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)Despite recent advances in image-to-video generation, better controllability and local animation are less explored.
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections