Browse State-of-the-Art › Video Generation
Video Generation
609 papers with code · 16 benchmarks · 24 datasets archive 2025-07-28
( Various Video Generation Tasks. Gif credit: MaGViT )
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
16 leaderboard tables shown for this task, 16 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 16 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
24 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 609 papers with code (1,466 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Jun 2017 71 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 1 pointer-only (licence)Generative Adversarial Networks (GANs) excel at creating realistic images with complex models for which maximum likelihood is infeasible.
-
2 Mar 2023 15 repositories listed Syntology ran 28 of 57 samples · 29 unverified · 8 pointer-only (licence)Through extensive experiments, we demonstrate that they outperform existing distillation techniques for diffusion models in one- and few-step sampling, achieving the new state-of-the-art FID of 3.
-
22 Aug 2018 14 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 1 pointer-only (licence)This paper presents a simple method for "do as I do" motion transfer: given a source video of a person dancing, we can transfer that performance to a novel (amateur) target after only a few minutes of the target subject…
-
23 Nov 2018 13 repositories listed Syntology ran 2 of 25 samples · 23 unverifiedAdditionally, we propose a first set of metrics to quantitatively evaluate the accuracy as well as the perceptual quality of the temporal evolution.
-
7 Apr 2022 5 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedGenerating temporally coherent high fidelity video is an important milestone in generative modeling research.
-
17 Jul 2017 5 repositories listedThe proposed framework generates a video by mapping a sequence of random vectors to a sequence of video frames.
-
28 Nov 2024 4 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We introduce Open-Sora Plan, an open-source project that aims to contribute a large generation model for generating desired high-resolution videos with long durations based on various user inputs.
-
5 Jan 2024 4 repositories listed Syntology ran 10 of 13 samples · 3 unverifiedWe propose a novel Latent Diffusion Transformer, namely Latte, for video generation.
-
18 Apr 2023 4 repositories listed Syntology ran 18 of 26 samples · 8 unverified · 1 pointer-only (licence)We first pre-train an LDM on images only; then, we turn the image generator into a video generator by introducing a temporal dimension to the latent space diffusion model and fine-tuning on encoded image sequences, i.
-
12 Jul 2022 4 repositories listedDrawing images of characters with desired poses is an essential but laborious task in anime production.
-
4 Apr 2018 4 repositories listed Syntology ran 4 of 16 samples · 12 unverified · 3 pointer-only (licence)However, learning to predict raw future observations, such as frames in a video, is exceedingly challenging -- the ambiguous nature of the problem can cause a naively designed model to average together possible futures…
-
21 Nov 2016 4 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedIn this paper, we propose a generative model, Temporal Generative Adversarial Nets (TGAN), which can learn a semantic representation of unlabeled videos, and is capable of generating videos.
-
14 Feb 2025 3 repositories listed Syntology ran 2 of 9 samples · 7 unverifiedWe present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length.
-
4 Jan 2025 3 repositories listedMultimodal Vision Language Models (VLMs) have emerged as a transformative topic at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through…
-
25 Dec 2024 3 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)However, on the one hand, aggressively reusing all the features cached in previous timesteps leads to a severe drop in generation quality.
-
1 Apr 2024 3 repositories listedFor instance, the widely-used CLIPScore measures the alignment between a (generated) image and text prompt, but it fails to produce reliable scores for complex prompts involving compositions of objects, attributes, and…
-
1 Dec 2023 3 repositories listed Syntology ran 13 of 23 samples · 10 unverifiedTo address these challenges, we introduce StyleCrafter, a generic method that enhances pre-trained T2V models with a style control adapter, enabling video generation in any style by providing a reference image.
-
25 Nov 2023 3 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedWe then explore the impact of finetuning our base model on high-quality data and train a text-to-video model that is competitive with closed-source video generation.
-
30 Oct 2023 3 repositories listed Syntology ran 7 of 12 samples · 5 unverifiedThe I2V model is designed to produce videos that strictly adhere to the content of the provided reference image, preserving its content, structure, and style.
-
23 Oct 2023 3 repositories listed Syntology ran 9 of 24 samples · 15 unverifiedWith the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress.
-
9 Oct 2023 3 repositories listed Syntology ran 12 of 20 samples · 8 unverifiedWhile Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation.
-
22 Dec 2022 3 repositories listed Syntology ran 3 of 6 samples · 3 unverifiedTo replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator.
-
9 Nov 2022 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)In light of this, we propose the Disentangled Objective Video Quality Evaluator (DOVER) to learn the quality of UGC videos based on the two perspectives.
-
29 Oct 2022 3 repositories listedIn this work, we evaluate the space learned by INR-V on diverse generative tasks such as video interpolation, novel video generation, video inversion, and video inpainting against the existing baselines.
-
20 Apr 2021 3 repositories listedWe present VideoGPT: a conceptually simple architecture for scaling likelihood based generative modeling to natural videos.
-
22 Jun 2020 3 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedWe consider the task of generating diverse and novel videos from a single video sample.
-
21 Oct 2019 3 repositories listedIn this paper, we focus on human motion transfer - generation of a video depicting a particular subject, observed in a single image, performing a series of motions exemplified by an auxiliary (driving) video.
-
5 Apr 2019 3 repositories listedWe introduce point-to-point video generation that controls the generation process with two control points: the targeted start- and end-frames.
-
21 Feb 2018 3 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Sample generations are both varied and sharp, even many frames into the future, and compare favorably to those from existing approaches.
-
27 Nov 2017 3 repositories listedFlowGAN generates optical flow, which contains only the edge and motion of the videos to be begerated.
Syntology lines on 21 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections