Browse State-of-the-Art › Text-to-Video Generation
Text-to-Video Generation
97 papers with code · 6 benchmarks · 14 datasets archive 2025-07-28
Ma grand-mère m’a raconté que quand elle était étudiante, elle avait un petit-ami. À l’âge de 18 ans, il a dû partir pour le service militaire, elle ne l’a pas attendu et elle a épousé quelqu’un d’autre. Quand ma grand-mère avait 58-59 ans, un homme (son premier amour) lui a envoyé une demande d’amis sur un réseau social, ils ont commencé à parler... En moins de six mois, ils ont décidé de se voir. Le trajet en train a duré deux jours et ils se sont finalement rencontrés. Cela fait maintenant deux ans qu’ils habitent ensemble et qu’ils nous rendent visite de temps en temps. Je réalise maintenant que leur amour l’un envers l’autre n’a jamais cessé.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| MSR-VTT (18 rows) | Snap Video (512x288) | Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis | — | — | Compare |
| UCF-101 (10 rows) | Snap Video (Zero-shot, 512x288) | Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis | — | — | Compare |
| EvalCrafter Text-to-Video (ECTV) Dataset (5 rows) | VideoCrafter2 | VideoCrafter2: Overcoming Data Limitations for High-Quality Video... | code | Syntology ran 14 of 21 samples · 7 unverified | Compare |
| Kinetics (1 row) | NUWA (128×128) | NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| Something-Something V2 (1 row) | MAGVIT | MAGVIT: Masked Generative Video Transformer | code | Syntology ran 1 of 10 samples · 9 unverified | Compare |
| WebVid (1 row) | VideoFactory | Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
14 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 97 papers with code (201 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Aug 2023 5 repositories listed Syntology ran 11 of 16 samples · 5 unverifiedThis paper introduces ModelScopeT2V, a text-to-video synthesis model that evolves from a text-to-image synthesis model (i.
-
5 Jan 2024 4 repositories listed Syntology ran 10 of 13 samples · 3 unverifiedWe propose a novel Latent Diffusion Transformer, namely Latte, for video generation.
-
3 Jun 2023 4 repositories listed Syntology ran 5 of 21 samples · 16 unverified · 1 pointer-only (licence)The pursuit of controllability as a higher standard of visual content creation has yielded remarkable progress in customizable image synthesis.
-
18 Apr 2023 4 repositories listed Syntology ran 18 of 26 samples · 8 unverified · 1 pointer-only (licence)We first pre-train an LDM on images only; then, we turn the image generator into a video generator by introducing a temporal dimension to the latent space diffusion model and fine-tuning on encoded image sequences, i.
-
1 Dec 2023 3 repositories listed Syntology ran 13 of 23 samples · 10 unverifiedTo address these challenges, we introduce StyleCrafter, a generic method that enhances pre-trained T2V models with a style control adapter, enabling video generation in any style by providing a reference image.
-
30 Oct 2023 3 repositories listed Syntology ran 7 of 12 samples · 5 unverifiedThe I2V model is designed to produce videos that strictly adhere to the content of the provided reference image, preserving its content, structure, and style.
-
22 Dec 2022 3 repositories listed Syntology ran 3 of 6 samples · 3 unverifiedTo replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator.
-
29 Dec 2024 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedTo facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content.
-
12 Aug 2024 2 repositories listed Syntology ran 7 of 7 samples · 0 unverifiedWe present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos aligned with text prompt, with a frame rate of 16 fps and resolution of…
-
26 Jun 2024 2 repositories listed Syntology ran 24 of 27 samples · 3 unverifiedWe propose a novel text-to-video (T2V) generation benchmark, ChronoMagic-Bench, to evaluate the temporal and metamorphic capabilities of the T2V models (e.
-
14 Jun 2024 2 repositories listedDespite their impressive capabilities, current Video-LMMs have not been evaluated for anomaly detection tasks, which is critical to their deployment in practical scenarios e.
-
8 Jun 2024 2 repositories listed Syntology ran 6 of 15 samples · 9 unverified · 15 pointer-only (licence)Motion-based controllable video generation offers the potential for creating captivating visual content.
-
7 Apr 2024 2 repositories listedRecent advances in Text-to-Video generation (T2V) have achieved remarkable success in synthesizing high-quality general videos from textual descriptions.
-
17 Jan 2024 2 repositories listed Syntology ran 14 of 21 samples · 7 unverified · 21 pointer-only (licence)Based on this stronger coupling, we shift the distribution to higher quality without motion degradation by finetuning spatial modules with high-quality images, resulting in a generic high-quality video model.
-
26 Sep 2023 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedTo this end, we propose LaVie, an integrated video generation framework that operates on cascaded video latent diffusion models, comprising a base T2V model, a temporal interpolation model, and a video super-resolution…
-
25 Sep 2023 2 repositories listed Syntology ran 5 of 11 samples · 6 unverifiedText-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt.
-
3 Apr 2023 2 repositories listed Syntology ran 2 of 10 samples · 8 unverifiedGenerating text-editable and pose-controllable character videos have an imperious demand in creating various digital human.
-
15 Mar 2023 2 repositories listedA diffusion probabilistic model (DPM), which constructs a forward diffusion process by gradually adding noise to data points and learns the reverse denoising process to generate new samples, has been shown to handle…
-
29 Sep 2022 2 repositories listedWe propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V).
-
17 Apr 2022 2 repositories listedAltogether, MUGEN can help progress research in many tasks in multimodal understanding and generation.
-
29 May 2025 1 repository listedVideo captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos.
-
23 May 2025 1 repository listedInstead of accumulating the entire generation history, the policy ranks and selects the top-K most contextually relevant tokens, allowing the model to maintain a fixed computational budget while preserving content…
-
17 May 2025 1 repository listedTo this end, we present AIGVE-60K, a comprehensive dataset and benchmark for AI-Generated Video Evaluation, which features (i) comprehensive tasks, encompassing 3, 050 extensive prompts across 20 fine-grained task…
-
31 Mar 2025 1 repository listedThese results show that On-device Sora enables efficient and high-quality video generation on resource-constrained mobile devices.
-
26 Mar 2025 1 repository listedFirst, we construct and refine a supervised fine-tuning (SFT) dataset based on principles of safety and alignment.
-
26 Mar 2025 1 repository listedScore-based or diffusion models generate high-quality tabular data, surpassing GAN-based and VAE-based models.
-
26 Mar 2025 1 repository listedExtensive experiments demonstrate that our video watermarking methods effectively protect video data by significantly reducing video annotation performance across various video-based LLMs, showcasing both stealthiness…
-
21 Mar 2025 1 repository listedDespite substantial progress in text-to-video generation, achieving precise and flexible control over fine-grained spatiotemporal attributes remains a significant unresolved challenge in video generation research.
-
3 Mar 2025 1 repository listedThe VideoUFO comprises over 1.
-
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation18 Feb 2025 1 repository listedThe training of controllable text-to-video (T2V) models relies heavily on the alignment between videos and captions, yet little existing research connects video caption evaluation with T2V generation assessment.
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections