Papers › Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked...

Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation

23 Nov 2022CVPR 2023 1arXiv:2211.12824archive 2025-07-28

Tsu-Jui Fu, Licheng Yu, Ning Zhang, Cheng-Yang Fu, Jong-Chyi Su, William Yang Wang, Sean Bell

Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head and tail is also crucial, but they have rarely been explored for video completion. Since there could be different outcomes from the hints of just a few frames, a system that can follow natural language to perform video completion may significantly improve controllability. Inspired by this, we introduce a novel task, text-guided video completion (TVC), which requests the model to generate a video from partial frames guided by an instruction. We then propose Multimodal Masked Video Generation (MMVG) to address this TVC task. During training, MMVG discretizes the video frames into visual tokens and masks most of them to perform video completion from any time point. At inference time, a single MMVG model can address all 3 cases of TVC, including video prediction, rewind, and infilling, by applying corresponding masking conditions. We evaluate MMVG in various video scenarios, including egocentric, animation, and gaming. Extensive experimental results indicate that MMVG is effective in generating high-quality visual appearances with text guidance for TVC.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

tsujuifu/pytorch_tvc officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Text-to-Video GenerationVideo GenerationVideo Prediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text-to-Video Generation MSR-VTT MMVG CLIPSIM 0.2644 #14 of 18 Archive leaderboard report
Text-to-Video Generation MSR-VTT MMVG FID 23.4 #14 of 18 Archive leaderboard report
Video Generation UCF-101 MMVG (128x128, class-conditional) FVD16 328 #24 of 48 Archive leaderboard report
Video Generation UCF-101 MMVG (128x128, class-conditional) Inception Score 73.7 #24 of 48 Archive leaderboard report
Video Generation UCF-101 MMVG (128x128, unconditional) FVD16 395 #33 of 48 Archive leaderboard report
Video Generation UCF-101 MMVG (128x128, unconditional) Inception Score 58.3 #33 of 48 Archive leaderboard report
Video Prediction BAIR Robot Pushing MMVG FVD 85.2 #3 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections