Papers › LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval

LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval

11 Jul 2022arXiv:2207.04858archive 2025-07-28

Jinbin Bai, Chunhui Liu, Feiyue Ni, Haofan Wang, Mengying Hu, Xiaofeng Guo, Lele Cheng

Video-text retrieval is a class of cross-modal representation learning problems, where the goal is to select the video which corresponds to the text query between a given text query and a pool of candidate videos. The contrastive paradigm of vision-language pretraining has shown promising success with large-scale datasets and unified transformer architecture, and demonstrated the power of a joint latent space. Despite this, the intrinsic divergence between the visual domain and textual domain is still far from being eliminated, and projecting different modalities into a joint latent space might result in the distorting of the information inside the single modality. To overcome the above issue, we present a novel mechanism for learning the translation relationship from a source modality space 𝒮 to a target modality space 𝒯 without the need for a joint latent space, which bridges the gap between visual and textual domains. Furthermore, to keep cycle consistency between translations, we adopt a cycle loss involving both forward translations from 𝒮 to the predicted target space 𝒯′, and backward translations from 𝒯′ back to 𝒮. Extensive experiments conducted on MSR-VTT, MSVD, and DiDeMo datasets demonstrate the superiority and effectiveness of our LaT approach compared with vanilla state-of-the-art methods.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Representation LearningRetrievalText RetrievalTranslationVideo RetrievalVideo-Text RetrievalZero-Shot Video Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Video Retrieval DiDeMo LaT text-to-video Median Rank 7 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo LaT text-to-video R@1 22.6 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo LaT text-to-video R@10 58.9 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo LaT text-to-video R@5 45.9 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo LaT video-to-text Median Rank 7 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo LaT video-to-text R@1 22.5 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo LaT video-to-text R@10 56.8 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo LaT video-to-text R@5 45.2 #23 of 26 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT text-to-video Median Rank 8 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT text-to-video R@1 23.4 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT text-to-video R@10 53.3 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT text-to-video R@5 44.1 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT video-to-text Median Rank 12 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT video-to-text R@1 17.2 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT video-to-text R@10 47.9 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT LaT video-to-text R@5 36.2 #32 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT text-to-video Median Rank 2 #13 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT text-to-video R@1 36.9 #13 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT text-to-video R@10 81.0 #13 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT text-to-video R@5 68.6 #13 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT video-to-text Median Rank 3 #13 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT video-to-text R@1 34.4 #13 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT video-to-text R@10 79.2 #13 of 14 Archive leaderboard report
Zero-Shot Video Retrieval MSVD LaT video-to-text R@5 69.0 #13 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections