{"url":"/task/image-to-video","name":"Image to Video Generation","slug":"image-to-video","description_markdown":"**Image to Video Generation** refers to the task of generating a sequence of video frames based on a single still image or a set of still images. The goal is to produce a video that is coherent and consistent in terms of appearance, motion, and style, while also being temporally consistent, meaning that the generated video should look like a coherent sequence of frames that are temporally ordered. This task is typically tackled using deep generative models, such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs), that are trained on large datasets of videos. The models learn to generate plausible video frames that are conditioned on the input image, as well as on any other auxiliary information, such as a sound or text track.","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":85,"papers_with_code":38,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":7,"subtasks":1,"parent_tasks":2},"benchmarks":[],"datasets":[{"url":"/dataset/opens2v-eval","name":"OpenS2V-Eval","full_name":"","num_papers_in_archive":8},{"url":"/dataset/chronomagic","name":"ChronoMagic","full_name":"","num_papers_in_archive":5},{"url":"/dataset/chronomagic-pro","name":"ChronoMagic-Pro","full_name":"","num_papers_in_archive":2},{"url":"/dataset/dropletvideo-10m","name":"DropletVideo-10M","full_name":"","num_papers_in_archive":2},{"url":"/dataset/chronomagic-proh","name":"ChronoMagic-ProH","full_name":"","num_papers_in_archive":1},{"url":"/dataset/consisid-preview-data","name":"ConsisID-preview-Data","full_name":"","num_papers_in_archive":1},{"url":"/dataset/opens2v-5m","name":"OpenS2V-5M","full_name":"","num_papers_in_archive":1}],"subtasks":[{"url":"/task/open-domain-subject-to-video","name":"Open-Domain Subject-to-Video"}],"parent_tasks":[{"url":"/task/1-image-2-2-stitchi","name":"1 Image, 2*2 Stitchi"},{"url":"/task/video-generation","name":"Video Generation"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":38,"tagged_in_all":85,"items":[{"url":"/paper/collaborative-neural-rendering-using-anime","title":"Collaborative Neural Rendering using Anime Character Sheets","date":"2022-07-12","arxiv_id":"2207.05378","repositories_listed":4,"syntology":null},{"url":"/paper/stable-video-diffusion-scaling-latent-video-1","title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","date":"2023-11-25","arxiv_id":"2311.15127","repositories_listed":3,"syntology":{"n":4,"n_ran":0,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/open-sora-democratizing-efficient-video","title":"Open-Sora: Democratizing Efficient Video Production for All","date":"2024-12-29","arxiv_id":"2412.20404","repositories_listed":2,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/scenerf-self-supervised-monocular-3d-scene","title":"SceneRF: Self-Supervised Monocular 3D Scene Reconstruction with Radiance Fields","date":"2022-12-05","arxiv_id":"2212.02501","repositories_listed":2,"syntology":null},{"url":"/paper/lifespan-age-transformation-synthesis","title":"Lifespan Age Transformation Synthesis","date":"2020-03-21","arxiv_id":"2003.09764","repositories_listed":2,"syntology":null},{"url":"/paper/video-generation-from-single-semantic-label","title":"Video Generation from Single Semantic Label Map","date":"2019-03-11","arxiv_id":"1903.04480","repositories_listed":2,"syntology":null},{"url":"/paper/every-painting-awakened-a-training-free","title":"Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation","date":"2025-03-31","arxiv_id":"2503.23736","repositories_listed":1,"syntology":null},{"url":"/paper/step-video-ti2v-technical-report-a-state-of","title":"Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model","date":"2025-03-14","arxiv_id":"2503.11251","repositories_listed":1,"syntology":null},{"url":"/paper/dualdiff-dual-branch-diffusion-for-high","title":"DualDiff+: Dual-Branch Diffusion for High-Fidelity Video Generation with Reward Guidance","date":"2025-03-05","arxiv_id":"2503.03689","repositories_listed":1,"syntology":null},{"url":"/paper/extrapolating-and-decoupling-image-to-video","title":"Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think","date":"2025-03-02","arxiv_id":"2503.00948","repositories_listed":1,"syntology":{"n":15,"n_ran":6,"n_unverified":9,"n_pointer_only":0}},{"url":"/paper/object-centric-image-to-video-generation-with","title":"Object-Centric Image to Video Generation with Language Guidance","date":"2025-02-17","arxiv_id":"2502.11655","repositories_listed":1,"syntology":null},{"url":"/paper/magic-1-for-1-generating-one-minute-video","title":"Magic 1-For-1: Generating One Minute Video Clips within One Minute","date":"2025-02-11","arxiv_id":"2502.07701","repositories_listed":1,"syntology":null},{"url":"/paper/framepainter-endowing-interactive-image","title":"FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors","date":"2025-01-14","arxiv_id":"2501.08225","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/ltx-video-realtime-video-latent-diffusion","title":"LTX-Video: Realtime Video Latent Diffusion","date":"2024-12-30","arxiv_id":"2501.00103","repositories_listed":1,"syntology":{"n":9,"n_ran":0,"n_unverified":9,"n_pointer_only":0}},{"url":"/paper/exploring-the-frontiers-of-animation-video","title":"AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era","date":"2024-12-13","arxiv_id":"2412.10255","repositories_listed":1,"syntology":null},{"url":"/paper/identity-preserving-text-to-video-generation","title":"Identity-Preserving Text-to-Video Generation by Frequency Decomposition","date":"2024-11-26","arxiv_id":"2411.17440","repositories_listed":1,"syntology":null},{"url":"/paper/vbench-comprehensive-and-versatile-benchmark","title":"VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models","date":"2024-11-20","arxiv_id":"2411.13503","repositories_listed":1,"syntology":null},{"url":"/paper/echomimicv2-towards-striking-simplified-and","title":"EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation","date":"2024-11-15","arxiv_id":"2411.10061","repositories_listed":1,"syntology":{"n":5,"n_ran":1,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/robust-watermarking-using-generative-priors","title":"Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances","date":"2024-10-24","arxiv_id":"2410.18775","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/redefining-temporal-modeling-in-video","title":"Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach","date":"2024-10-04","arxiv_id":"2410.03160","repositories_listed":1,"syntology":{"n":10,"n_ran":7,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/physgen-rigid-body-physics-grounded-image-to","title":"PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation","date":"2024-09-27","arxiv_id":"2409.18964","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/genrec-unifying-video-generation-and","title":"GenRec: Unifying Video Generation and Recognition with Diffusion Models","date":"2024-08-27","arxiv_id":"2408.15241","repositories_listed":1,"syntology":null},{"url":"/paper/mmtrail-a-multimodal-trailer-video-dataset","title":"MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions","date":"2024-07-30","arxiv_id":"2407.20962","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/mvoc-a-training-free-multiple-video-object","title":"MVOC: a training-free multiple video object composition method with diffusion models","date":"2024-06-22","arxiv_id":"2406.15829","repositories_listed":1,"syntology":null},{"url":"/paper/tc-bench-benchmarking-temporal","title":"TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation","date":"2024-06-12","arxiv_id":"2406.08656","repositories_listed":1,"syntology":{"n":5,"n_ran":5,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/ti2v-zero-zero-shot-image-conditioning-for","title":"TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models","date":"2024-04-25","arxiv_id":"2404.16306","repositories_listed":1,"syntology":null},{"url":"/paper/anyv2v-a-plug-and-play-framework-for-any","title":"AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks","date":"2024-03-21","arxiv_id":"2403.14468","repositories_listed":1,"syntology":{"n":7,"n_ran":2,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/champ-controllable-and-consistent-human-image","title":"Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance","date":"2024-03-21","arxiv_id":"2403.14781","repositories_listed":1,"syntology":{"n":5,"n_ran":1,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/mora-enabling-generalist-video-generation-via","title":"Mora: Enabling Generalist Video Generation via A Multi-Agent Framework","date":"2024-03-20","arxiv_id":"2403.13248","repositories_listed":1,"syntology":{"n":4,"n_ran":4,"n_unverified":0,"n_pointer_only":4}},{"url":"/paper/follow-your-click-open-domain-regional-image","title":"Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts","date":"2024-03-13","arxiv_id":"2403.08268","repositories_listed":1,"syntology":{"n":9,"n_ran":5,"n_unverified":4,"n_pointer_only":9}}],"syntology_records":15,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}