Browse State-of-the-Art › Video Inpainting
Video Inpainting
59 papers with code · 7 benchmarks · 15 datasets archive 2025-07-28
The goal of Video Inpainting is to fill in missing regions of a given video sequence with contents that are both spatially and temporally coherent. Video Inpainting, also known as video completion, has many real-world applications such as undesired object removal and video restoration.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| DAVIS (11 rows) | DMT | Deficiency-Aware Masked Transformer for Video Inpainting | code | — | Compare |
| YouTube-VOS 2018 (10 rows) | ProPainter | ProPainter: Improving Propagation and Transformer for Video Inpainting | code | Syntology ran 20 of 34 samples · 14 unverified | Compare |
| HQVI (240p) (7 rows) | RGVI | Elevating Flow-Guided Video Inpainting with Reference Generation | code | — | Compare |
| HQVI (480p) (4 rows) | RGVI | Elevating Flow-Guided Video Inpainting with Reference Generation | code | — | Compare |
| HQVI (2K) (2 rows) | RGVI | Elevating Flow-Guided Video Inpainting with Reference Generation | code | — | Compare |
| YouTube-VOS (2 rows) | FGT++ | Exploiting Optical Flow Guidance for Transformer-Based Video Inpainting | code | — | Compare |
| How2Sign (1 row) | INR-V | INR-V: A Continuous Representation Space for Video-based Generative Tasks | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
15 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 59 papers with code (130 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
7 Sep 2023 3 repositories listed Syntology ran 20 of 34 samples · 14 unverified · 16 pointer-only (licence)We also propose a mask-guided sparse video Transformer, which achieves high efficiency by discarding unnecessary and redundant tokens.
-
21 Nov 2022 3 repositories listedIn this paper, we propose the concept of online video inpainting for autonomous vehicles to expand the field of view, thereby enhancing scene visibility, perception, and system safety.
-
29 Oct 2022 3 repositories listedIn this work, we evaluate the space learned by INR-V on diverse generative tasks such as video interpolation, novel video generation, video inversion, and video inpainting against the existing baselines.
-
24 Jan 2023 2 repositories listedTransformers have been widely used for video processing owing to the multi-head self attention (MHSA) mechanism.
-
6 Apr 2022 2 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)Optical flow, which captures motion information across frames, is exploited in recent video inpainting methods through propagating pixels along its trajectories.
-
20 Jul 2020 2 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedIn this paper, we propose to learn a joint Spatial-Temporal Transformer Network (STTN) for video inpainting.
-
17 Jul 2020 2 repositories listed Syntology ran 2 of 10 samples · 8 unverifiedTo get clear street-view and photo-realistic simulation in autonomous driving, we present an automatic video inpainting algorithm that can remove traffic agents from videos and synthesize missing regions with the…
-
2 Jul 2019 2 repositories listedHow to efficiently utilize temporal information to recover videos in a consistent way is the main issue for video inpainting problems.
-
8 May 2019 2 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedThen the synthesized flow field is used to guide the propagation of pixels to fill up the missing regions in the video.
-
5 May 2019 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Video inpainting aims to fill spatio-temporal holes with plausible content in a video.
-
23 Apr 2019 2 repositories listedFree-form video inpainting is a very challenging task that could be widely used for video editing such as text removal.
-
12 Mar 2025 1 repository listedVideo body-swapping aims to replace the body in an existing video with a new body from arbitrary sources, which has garnered more attention in recent years.
-
7 Mar 2025 1 repository listedVideo inpainting, which aims to restore corrupted video content, has experienced substantial progress.
-
17 Jan 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedRecent video inpainting algorithms integrate flow-based pixel propagation with transformer-based generation to leverage optical flow for restoring textures and objects using information from neighboring frames, while…
-
12 Dec 2024 1 repository listedPowered by a strong generative model, our method not only significantly enhances frame-level quality for object removal but also synthesizes new content in the missing areas based on user-provided text prompts.
-
1 Dec 2024 1 repository listedSpecifically, FloED employs a dual-branch architecture, where a flow branch first restores corrupted flow and a multi-scale flow adapter provides motion guidance to the main inpainting branch.
-
10 Oct 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedThe spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting.
-
30 Sep 2024 1 repository listedFurthermore, we implement an enriched visual guidance mechanism to enhance appearance alignment, a hybrid inpainting encoder to further preserve the detailed background information in the masked video, and a two-phase…
-
11 Sep 2024 1 repository listedThis paper presents a novel framework for converting 2D videos to immersive stereoscopic 3D, addressing the growing demand for 3D content in immersive experience.
-
2 Jul 2024 1 repository listedVideo inpainting fills in corrupted video content with plausible replacements.
-
25 Jun 2024 1 repository listedThis letter proposes a simple yet effective forensic scheme for Video Inpainting LOcalization with ContrAstive Learning (ViLocal).
-
19 Jun 2024 1 repository listedIn this paper, we present a Trusted Video Inpainting Localization network (TruVIL) with excellent robustness and generalization ability.
-
24 Apr 2024 1 repository listedVideo Wire Inpainting (VWI) is a prominent application in video inpainting, aimed at flawlessly removing wires in films or TV series, offering significant time and labor savings compared to manual frame-by-frame removal.
-
19 Mar 2024 1 repository listedThis block incorporates a codebook mechanism to discretize the network's shallow residual features and inter-frame residual information effectively.
-
18 Mar 2024 1 repository listedTo this end, this paper proposes a novel text-guided video inpainting model that achieves better consistency, controllability and compatibility.
-
25 Jan 2024 1 repository listedOur results demonstrate the remarkable capability of the proposed framework to remove HMDs from facial videos while maintaining the subject's facial expression and identity.
-
23 Jan 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedWe introduce Lumiere -- a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion -- a pivotal challenge in video synthesis.
-
18 Jan 2024 1 repository listed Syntology ran 7 of 10 samples · 3 unverified · 10 pointer-only (licence)We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process.
-
6 Dec 2023 1 repository listedGiven a video, a masked region at its initial frame, and an editing prompt, it requires a model to do infilling at each frame following the editing guidance while keeping the out-of-mask region intact.
-
5 Dec 2023 1 repository listed Syntology ran 3 of 6 samples · 3 unverifiedNow text-to-image foundation models are widely applied to various downstream image synthesis tasks, such as controllable image generation and image editing, while downstream video synthesis tasks are less explored for…
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections