Papers › MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

26 Apr 2022arXiv:2204.12408archive 2025-07-28

Yuying Ge, Yixiao Ge, Xihui Liu, Alex Jinpeng Wang, Jianping Wu, Ying Shan, XiaoHu Qie, Ping Luo

Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast global video and text representations, but ignore detailed local semantics. The recent success of image BERT pre-training with masked visual modeling that promotes the learning of local visual context, motivates a possible solution to address the above limitation. In this work, we for the first time investigate masked visual modeling in video-text pre-training with the "dual-encoder" architecture. We perform Masked visual modeling with Injected LanguagE Semantics (MILES) by employing an extra snapshot video encoder as an evolving "tokenizer" to produce reconstruction targets for masked video patch prediction. Given the corrupted video, the video encoder is trained to recover text-aligned features of the masked patches via reasoning with the visible regions along the spatial and temporal dimensions, which enhances the discriminativeness of local visual features and the fine-grained cross-modality alignment. Our method outperforms state-of-the-art methods for text-to-video retrieval on four datasets with both zero-shot and fine-tune evaluation protocols. Our approach also surpasses the baseline models significantly on zero-shot action recognition, which can be cast as video-to-text retrieval.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

tencentarc/mcq officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionRetrievalText RetrievalText to Video RetrievalVideo RetrievalVideo to Text RetrievalVideo-Text RetrievalZero-Shot Action RecognitionZero-Shot Video Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Video Retrieval DiDeMo MILES text-to-video Median Rank 5.0 #18 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo MILES text-to-video R@1 27.2 #18 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo MILES text-to-video R@10 63.6 #18 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo MILES text-to-video R@5 50.3 #18 of 26 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC MILES text-to-video Median Rank 50.7 #15 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC MILES text-to-video R@1 11.1 #15 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC MILES text-to-video R@10 30.6 #15 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC MILES text-to-video R@5 24.7 #15 of 16 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT MILES text-to-video Median Rank 7 #26 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT MILES text-to-video R@1 26.1 #26 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT MILES text-to-video R@10 56.9 #26 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT MILES text-to-video R@5 47.2 #26 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSVD MILES text-to-video Median Rank 2.0 #9 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD MILES text-to-video R@1 44.4 #9 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD MILES text-to-video R@10 87.0 #9 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD MILES text-to-video R@5 76.2 #9 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections