{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/align-and-prompt-video-and-language-pre","title":"Align and Prompt: Video-and-Language Pre-training with Entity Prompts","arxiv_id":"2112.09583","date":"2021-12-17","proceeding":"CVPR 2022 1","authors":["Dongxu Li","Junnan Li","Hongdong Li","Juan Carlos Niebles","Steven C. H. Hoi"],"abstract":"Video-and-language pre-training has shown promising improvements on various downstream tasks. Most previous methods capture cross-modal interactions with a transformer-based multimodal encoder, not fully addressing the misalignment between unimodal video and text features. Besides, learning fine-grained visual-language alignment usually requires off-the-shelf object detectors to provide object information, which is bottlenecked by the detector's limited vocabulary and expensive computation cost. We propose Align and Prompt: an efficient and effective video-and-language pre-training framework with better cross-modal alignment. First, we introduce a video-text contrastive (VTC) loss to align unimodal video-text features at the instance level, which eases the modeling of cross-modal interactions. Then, we propose a new visually-grounded pre-training task, prompting entity modeling (PEM), which aims to learn fine-grained region-entity alignment. To achieve this, we first introduce an entity prompter module, which is trained with VTC to produce the similarity between a video crop and text prompts instantiated with entity names. The PEM task then asks the model to predict the entity pseudo-labels (i.e~normalized similarity scores) for randomly-selected video crops. The resulting pre-trained model achieves state-of-the-art performance on both text-video retrieval and videoQA, outperforming prior work by a substantial margin. Our code and pre-trained models are available at https://github.com/salesforce/ALPRO.","url_abs":"https://arxiv.org/abs/2112.09583v2","url_pdf":"https://arxiv.org/pdf/2112.09583v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"align-and-prompt-video-and-language-pre","repo_url":"https://github.com/salesforce/alpro","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"entity-alignment","task_name":"Entity Alignment"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"zero-shot-video-retrieval","task_name":"Zero-Shot Video Retrieval"},{"task_slug":"cross-modal-alignment","task_name":"cross-modal alignment"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-retrieval-on-didemo","task":"Video Retrieval","dataset":"DiDeMo","model":"ALPRO","rank_in_archive_order":35,"of":40,"metrics":{"text-to-video Median Rank":"3","text-to-video R@1":"35.9","text-to-video R@10":"78.8","text-to-video R@5":"67.5"},"uses_additional_data":true},{"leaderboard":"/sota/visual-question-answering-on-msrvtt-qa-1","task":"Visual Question Answering (VQA)","dataset":"MSRVTT-QA","model":"ALPRO","rank_in_archive_order":23,"of":34,"metrics":{"Accuracy":"0.421"},"uses_additional_data":true},{"leaderboard":"/sota/visual-question-answering-on-msvd-qa-1","task":"Visual Question Answering (VQA)","dataset":"MSVD-QA","model":"ALPRO","rank_in_archive_order":29,"of":36,"metrics":{"Accuracy":"0.459"},"uses_additional_data":true},{"leaderboard":"/sota/zero-shot-video-retrieval-on-didemo","task":"Zero-Shot Video Retrieval","dataset":"DiDeMo","model":"ALPRO","rank_in_archive_order":20,"of":26,"metrics":{"text-to-video Median Rank":"6","text-to-video R@1":"23.8","text-to-video R@10":"57.9","text-to-video R@5":"47.3"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-retrieval-on-msr-vtt","task":"Zero-Shot Video Retrieval","dataset":"MSR-VTT","model":"ALPRO","rank_in_archive_order":30,"of":41,"metrics":{"text-to-video Median Rank":"8","text-to-video R@1":"24.1","text-to-video R@10":"55.4","text-to-video R@5":"44.7"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2112.09583","atlas_url":"https://app.syntology.ai/?focus=2112.09583","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}