Papers › Deficiency-Aware Masked Transformer for Video Inpainting

Deficiency-Aware Masked Transformer for Video Inpainting

17 Jul 2023arXiv:2307.08629archive 2025-07-28

Yongsheng Yu, Heng Fan, Libo Zhang

Recent video inpainting methods have made remarkable progress by utilizing explicit guidance, such as optical flow, to propagate cross-frame pixels. However, there are cases where cross-frame recurrence of the masked video is not available, resulting in a deficiency. In such situation, instead of borrowing pixels from other frames, the focus of the model shifts towards addressing the inverse problem. In this paper, we introduce a dual-modality-compatible inpainting framework called Deficiency-aware Masked Transformer (DMT), which offers three key advantages. Firstly, we pretrain a image inpainting model DMT_img serve as a prior for distilling the video model DMT_vid, thereby benefiting the hallucination of deficiency cases. Secondly, the self-attention module selectively incorporates spatiotemporal tokens to accelerate inference and remove noise signals. Thirdly, a simple yet effective Receptive Field Contextualizer is integrated into DMT, further improving performance. Extensive experiments conducted on YouTube-VOS and DAVIS datasets demonstrate that DMT_vid significantly outperforms previous solutions. The code and video demonstrations can be found at github.com/yeates/DMT.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

yeates/dmt officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

HallucinationImage InpaintingOptical Flow EstimationVideo Inpainting

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Inpainting DAVIS DMT Ewarp - #1 of 11 Archive leaderboard report
Video Inpainting DAVIS DMT PSNR 33.82 #1 of 11 Archive leaderboard report
Video Inpainting DAVIS DMT SSIM 0.976 #1 of 11 Archive leaderboard report
Video Inpainting DAVIS DMT VFID 0.104 #1 of 11 Archive leaderboard report
Video Inpainting YouTube-VOS 2018 DMT Ewarp - #2 of 10 Archive leaderboard report
Video Inpainting YouTube-VOS 2018 DMT PSNR 34.27 #2 of 10 Archive leaderboard report
Video Inpainting YouTube-VOS 2018 DMT SSIM 0.9730 #2 of 10 Archive leaderboard report
Video Inpainting YouTube-VOS 2018 DMT VFID 0.044 #2 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutFocusInpaintingLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections