Papers › An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling

An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling

4 Sep 2022CVPR 2023 1arXiv:2209.01540archive 2025-07-28

Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, Zicheng Liu

Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) pre-training, previous studies fail to find a truly effective MVM strategy that can largely benefit the downstream performance. In this work, we systematically examine the potential of MVM in the context of VidL learning. Specifically, we base our study on a fully end-to-end VIdeO-LanguagE Transformer (VIOLET), where the supervision from MVM training can be backpropagated to the video pixel space. In total, eight different reconstructive targets of MVM are explored, from low-level pixel values and oriented gradients to high-level depth maps, optical flow, discrete visual tokens, and latent visual features. We conduct comprehensive experiments and provide insights into the factors leading to effective MVM training, resulting in an enhanced model VIOLETv2. Empirically, we show VIOLETv2 pre-trained with MVM objective achieves notable improvements on 13 VidL benchmarks, ranging from video question answering, video captioning, to text-to-video retrieval.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

tsujuifu/pytorch_empirical-mvm officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Fill MaskOptical Flow EstimationQuestion AnsweringRetrievalTGIF-ActionTGIF-FrameTGIF-TransitionText to Video RetrievalVideo CaptioningVideo Question AnsweringVideo RetrievalVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Captioning MSR-VTT VIOLETv2 CIDEr 58 #18 of 24 Archive leaderboard report
Video Captioning MSVD VIOLETv2 CIDEr 139.2 #9 of 16 Archive leaderboard report
Video Question Answering LSMDC-MC VIOLETv2 Accuracy 84.4 #1 of 2 Archive leaderboard report
Video Question Answering MSRVTT-MC VIOLETv2 Accuracy 97.6 #1 of 7 Archive leaderboard report
Video Question Answering MSRVTT-QA VIOLETv2 Accuracy 44.5 #11 of 14 Archive leaderboard report
Video Retrieval DiDeMo VIOLETv2 text-to-video R@1 47.9 #28 of 40 Archive leaderboard report
Video Retrieval DiDeMo VIOLETv2 text-to-video R@10 84.1 #28 of 40 Archive leaderboard report
Video Retrieval DiDeMo VIOLETv2 text-to-video R@5 76.5 #28 of 40 Archive leaderboard report
Video Retrieval LSMDC VIOLETv2 text-to-video R@1 24 #21 of 38 Archive leaderboard report
Video Retrieval LSMDC VIOLETv2 text-to-video R@10 54.1 #21 of 38 Archive leaderboard report
Video Retrieval LSMDC VIOLETv2 text-to-video R@5 43.5 #21 of 38 Archive leaderboard report
Video Retrieval MSR-VTT VIOLETv2 text-to-video R@1 37.2 #16 of 40 Archive leaderboard report
Video Retrieval MSR-VTT VIOLETv2 text-to-video R@10 75.8 #16 of 40 Archive leaderboard report
Video Retrieval MSR-VTT VIOLETv2 text-to-video R@5 64.8 #16 of 40 Archive leaderboard report
Visual Question Answering (VQA) MSVD-QA VIOLETv2 Accuracy 0.547 #15 of 36 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBASEBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections