{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lat-latent-translation-with-cycle-consistency","title":"LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval","arxiv_id":"2207.04858","date":"2022-07-11","proceeding":null,"authors":["Jinbin Bai","Chunhui Liu","Feiyue Ni","Haofan Wang","Mengying Hu","Xiaofeng Guo","Lele Cheng"],"abstract":"Video-text retrieval is a class of cross-modal representation learning problems, where the goal is to select the video which corresponds to the text query between a given text query and a pool of candidate videos. The contrastive paradigm of vision-language pretraining has shown promising success with large-scale datasets and unified transformer architecture, and demonstrated the power of a joint latent space. Despite this, the intrinsic divergence between the visual domain and textual domain is still far from being eliminated, and projecting different modalities into a joint latent space might result in the distorting of the information inside the single modality. To overcome the above issue, we present a novel mechanism for learning the translation relationship from a source modality space $\\mathcal{S}$ to a target modality space $\\mathcal{T}$ without the need for a joint latent space, which bridges the gap between visual and textual domains. Furthermore, to keep cycle consistency between translations, we adopt a cycle loss involving both forward translations from $\\mathcal{S}$ to the predicted target space $\\mathcal{T'}$, and backward translations from $\\mathcal{T'}$ back to $\\mathcal{S}$. Extensive experiments conducted on MSR-VTT, MSVD, and DiDeMo datasets demonstrate the superiority and effectiveness of our LaT approach compared with vanilla state-of-the-art methods.","url_abs":"https://arxiv.org/abs/2207.04858v2","url_pdf":"https://arxiv.org/pdf/2207.04858v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"},{"task_slug":"video-text-retrieval","task_name":"Video-Text Retrieval"},{"task_slug":"zero-shot-video-retrieval","task_name":"Zero-Shot Video Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-video-retrieval-on-didemo","task":"Zero-Shot Video Retrieval","dataset":"DiDeMo","model":"LaT","rank_in_archive_order":23,"of":26,"metrics":{"text-to-video Median Rank":"7","text-to-video R@1":"22.6","text-to-video R@10":"58.9","text-to-video R@5":"45.9","video-to-text Median Rank":"7","video-to-text R@1":"22.5","video-to-text R@10":"56.8","video-to-text R@5":"45.2"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-retrieval-on-msr-vtt","task":"Zero-Shot Video Retrieval","dataset":"MSR-VTT","model":"LaT","rank_in_archive_order":32,"of":41,"metrics":{"text-to-video Median Rank":"8","text-to-video R@1":"23.4","text-to-video R@10":"53.3","text-to-video R@5":"44.1","video-to-text Median Rank":"12","video-to-text R@1":"17.2","video-to-text R@10":"47.9","video-to-text R@5":"36.2"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-retrieval-on-msvd","task":"Zero-Shot Video Retrieval","dataset":"MSVD","model":"LaT","rank_in_archive_order":13,"of":14,"metrics":{"text-to-video Median Rank":"2","text-to-video R@1":"36.9","text-to-video R@10":"81.0","text-to-video R@5":"68.6","video-to-text Median Rank":"3","video-to-text R@1":"34.4","video-to-text R@10":"79.2","video-to-text R@5":"69.0"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}