{"url":"/task/zero-shot-video-retrieval","name":"Zero-Shot Video Retrieval","slug":"zero-shot-video-retrieval","description_markdown":"Zero-shot video retrieval is the task of retrieving relevant videos based on a query (usually in text form) without any prior training on specific examples of those videos. Unlike traditional retrieval methods that rely on supervised learning with annotated datasets, zero-shot retrieval leverages pre-trained models, typically based on large-scale vision-language learning, to understand semantic relationships between textual descriptions and video content.\r\n\r\nThis approach enables retrieval of unseen video concepts by generalizing knowledge from diverse training data, making it highly useful for domains with limited labeled data, such as broadcast media, surveillance, and historical archives.","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":40,"papers_with_code":33,"benchmarks":8,"benchmark_tables_in_archive":8,"benchmark_tables_shown":8,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":7,"subtasks":0,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/zero-shot-video-retrieval-on-msr-vtt","slug":"zero-shot-video-retrieval-on-msr-vtt","dataset":"MSR-VTT","dataset_url":"/dataset/msr-vtt","rows_in_archive":41,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","text-to-video Median Rank","text-to-video Mean Rank","video-to-text R@1","video-to-text R@5","video-to-text R@10","video-to-text Median Rank"],"first_row_in_archive_order":{"model":"InternVideo2-6B","paper_title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","paper_url":"/paper/internvideo2-scaling-video-foundation-models","paper_date":"2024-03-22","arxiv_id":"2403.15377","code_links":[{"title":"opengvlab/internvideo","url":"https://github.com/opengvlab/internvideo"},{"title":"opengvlab/internvideo2","url":"https://github.com/opengvlab/internvideo2"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-video-retrieval-on-didemo","slug":"zero-shot-video-retrieval-on-didemo","dataset":"DiDeMo","dataset_url":"/dataset/didemo","rows_in_archive":26,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","video-to-text R@1","video-to-text R@5","video-to-text R@10","text-to-video Median Rank","video-to-text Median Rank"],"first_row_in_archive_order":{"model":"InternVideo2-6B","paper_title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","paper_url":"/paper/internvideo2-scaling-video-foundation-models","paper_date":"2024-03-22","arxiv_id":"2403.15377","code_links":[{"title":"opengvlab/internvideo","url":"https://github.com/opengvlab/internvideo"},{"title":"opengvlab/internvideo2","url":"https://github.com/opengvlab/internvideo2"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-video-retrieval-on-lsmdc","slug":"zero-shot-video-retrieval-on-lsmdc","dataset":"LSMDC","dataset_url":"/dataset/lsmdc","rows_in_archive":16,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","text-to-video Median Rank","text-to-video Mean Rank","video-to-text R@1","video-to-text R@5","video-to-text R@10"],"first_row_in_archive_order":{"model":"InternVideo2-6B","paper_title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","paper_url":"/paper/internvideo2-scaling-video-foundation-models","paper_date":"2024-03-22","arxiv_id":"2403.15377","code_links":[{"title":"opengvlab/internvideo","url":"https://github.com/opengvlab/internvideo"},{"title":"opengvlab/internvideo2","url":"https://github.com/opengvlab/internvideo2"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-video-retrieval-on-msvd","slug":"zero-shot-video-retrieval-on-msvd","dataset":"MSVD","dataset_url":"/dataset/msvd","rows_in_archive":14,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","text-to-video Median Rank","text-to-video Mean Rank","video-to-text R@1","video-to-text R@5","video-to-text R@10","video-to-text Median Rank"],"first_row_in_archive_order":{"model":"InternVideo2-6B","paper_title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","paper_url":"/paper/internvideo2-scaling-video-foundation-models","paper_date":"2024-03-22","arxiv_id":"2403.15377","code_links":[{"title":"opengvlab/internvideo","url":"https://github.com/opengvlab/internvideo"},{"title":"opengvlab/internvideo2","url":"https://github.com/opengvlab/internvideo2"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-video-retrieval-on-activitynet","slug":"zero-shot-video-retrieval-on-activitynet","dataset":"ActivityNet","dataset_url":"/dataset/activitynet","rows_in_archive":12,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","video-to-text R@1","video-to-text R@5","video-to-text R@10"],"first_row_in_archive_order":{"model":"InternVideo2-6B","paper_title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","paper_url":"/paper/internvideo2-scaling-video-foundation-models","paper_date":"2024-03-22","arxiv_id":"2403.15377","code_links":[{"title":"opengvlab/internvideo","url":"https://github.com/opengvlab/internvideo"},{"title":"opengvlab/internvideo2","url":"https://github.com/opengvlab/internvideo2"}],"syntology":null}},{"leaderboard":"/sota/zero-shot-video-retrieval-on-youcook2","slug":"zero-shot-video-retrieval-on-youcook2","dataset":"YouCook2","dataset_url":"/dataset/youcook2","rows_in_archive":9,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","text-to-video Mean Rank","text-to-video Median Rank"],"first_row_in_archive_order":{"model":"OmniVec2","paper_title":"OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning","paper_url":"/paper/omnivec2-a-novel-transformer-based-network","paper_date":"2024-01-01","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/zero-shot-video-retrieval-on-vatex","slug":"zero-shot-video-retrieval-on-vatex","dataset":"VATEX","dataset_url":"/dataset/vatex","rows_in_archive":5,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","video-to-text R@1","video-to-text R@5","video-to-text R@10"],"first_row_in_archive_order":{"model":"GRAM","paper_title":"Gramian Multimodal Representation Learning and Alignment","paper_url":"/paper/gramian-multimodal-representation-learning","paper_date":"2024-12-16","arxiv_id":"2412.11959","code_links":[{"title":"ispamm/GRAM","url":"https://github.com/ispamm/GRAM"},{"title":"luigisigillo/gwit","url":"https://github.com/luigisigillo/gwit"}],"syntology":{"n":12,"n_ran":2,"n_unverified":10,"n_pointer_only":0}}},{"leaderboard":"/sota/zero-shot-video-retrieval-on-msr-vtt-full","slug":"zero-shot-video-retrieval-on-msr-vtt-full","dataset":"MSR-VTT-full","dataset_url":null,"rows_in_archive":3,"metrics":["text-to-video R@1","text-to-video R@5","text-to-video R@10","video-to-text R@1","video-to-text R@5","video-to-text R@10"],"first_row_in_archive_order":{"model":"InternVL-G","paper_title":"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks","paper_url":"/paper/internvl-scaling-up-vision-foundation-models","paper_date":"2023-12-21","arxiv_id":"2312.14238","code_links":[{"title":"opengvlab/internvl","url":"https://github.com/opengvlab/internvl"},{"title":"opengvlab/internvl-mmdetseg","url":"https://github.com/opengvlab/internvl-mmdetseg"}],"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}}}],"datasets":[{"url":"/dataset/activitynet","name":"ActivityNet","full_name":"","num_papers_in_archive":807},{"url":"/dataset/msr-vtt","name":"MSR-VTT","full_name":"","num_papers_in_archive":640},{"url":"/dataset/msvd","name":"MSVD","full_name":"Microsoft Research Video Description Corpus","num_papers_in_archive":327},{"url":"/dataset/didemo","name":"DiDeMo","full_name":"Distinct Describable Moments","num_papers_in_archive":216},{"url":"/dataset/youcook2","name":"YouCook2","full_name":"","num_papers_in_archive":198},{"url":"/dataset/lsmdc","name":"LSMDC","full_name":"Large Scale Movie Description Challenge","num_papers_in_archive":126},{"url":"/dataset/vatex","name":"VATEX","full_name":"Video And TEXt","num_papers_in_archive":118}],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":33,"tagged_in_all":40,"items":[{"url":"/paper/languagebind-extending-video-language","title":"LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment","date":"2023-10-03","arxiv_id":"2310.01852","repositories_listed":6,"syntology":{"n":14,"n_ran":7,"n_unverified":7,"n_pointer_only":0}},{"url":"/paper/vatt-transformers-for-multimodal-self","title":"VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text","date":"2021-04-22","arxiv_id":"2104.11178","repositories_listed":5,"syntology":{"n":8,"n_ran":5,"n_unverified":3,"n_pointer_only":8}},{"url":"/paper/clip4clip-an-empirical-study-of-clip-for-end","title":"CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval","date":"2021-04-18","arxiv_id":"2104.08860","repositories_listed":5,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/frozen-in-time-a-joint-video-and-image","title":"Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval","date":"2021-04-01","arxiv_id":"2104.00650","repositories_listed":5,"syntology":{"n":11,"n_ran":3,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/mplug-2-a-modularized-multi-modal-foundation","title":"mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video","date":"2023-02-01","arxiv_id":"2302.00402","repositories_listed":4,"syntology":{"n":19,"n_ran":9,"n_unverified":10,"n_pointer_only":0}},{"url":"/paper/end-to-end-learning-of-visual-representations","title":"End-to-End Learning of Visual Representations from Uncurated Instructional Videos","date":"2019-12-13","arxiv_id":"1912.06430","repositories_listed":4,"syntology":{"n":5,"n_ran":3,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/imagebind-one-embedding-space-to-bind-them","title":"ImageBind: One Embedding Space To Bind Them All","date":"2023-05-09","arxiv_id":"2305.05665","repositories_listed":3,"syntology":{"n":34,"n_ran":24,"n_unverified":10,"n_pointer_only":32}},{"url":"/paper/gramian-multimodal-representation-learning","title":"Gramian Multimodal Representation Learning and Alignment","date":"2024-12-16","arxiv_id":"2412.11959","repositories_listed":2,"syntology":{"n":12,"n_ran":2,"n_unverified":10,"n_pointer_only":0}},{"url":"/paper/internvideo2-scaling-video-foundation-models","title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","date":"2024-03-22","arxiv_id":"2403.15377","repositories_listed":2,"syntology":null},{"url":"/paper/internvl-scaling-up-vision-foundation-models","title":"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks","date":"2023-12-21","arxiv_id":"2312.14238","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/vast-a-vision-audio-subtitle-text-omni-1","title":"VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset","date":"2023-05-29","arxiv_id":"2305.18500","repositories_listed":2,"syntology":{"n":42,"n_ran":15,"n_unverified":27,"n_pointer_only":0}},{"url":"/paper/internvideo-general-video-foundation-models","title":"InternVideo: General Video Foundation Models via Generative and Discriminative Learning","date":"2022-12-06","arxiv_id":"2212.03191","repositories_listed":2,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/revealing-single-frame-bias-for-video-and","title":"Revealing Single Frame Bias for Video-and-Language Learning","date":"2022-06-07","arxiv_id":"2206.03428","repositories_listed":2,"syntology":{"n":12,"n_ran":4,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/bridgeformer-bridging-video-text-retrieval","title":"Bridging Video-text Retrieval with Multiple Choice Questions","date":"2022-01-13","arxiv_id":"2201.04850","repositories_listed":2,"syntology":{"n":24,"n_ran":13,"n_unverified":11,"n_pointer_only":6}},{"url":"/paper/florence-a-new-foundation-model-for-computer","title":"Florence: A New Foundation Model for Computer Vision","date":"2021-11-22","arxiv_id":"2111.11432","repositories_listed":2,"syntology":null},{"url":"/paper/videoclip-contrastive-pre-training-for-zero","title":"VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding","date":"2021-09-28","arxiv_id":"2109.14084","repositories_listed":2,"syntology":null},{"url":"/paper/make-your-training-flexible-towards","title":"Make Your Training Flexible: Towards Deployment-Efficient Video Models","date":"2025-03-18","arxiv_id":"2503.14237","repositories_listed":1,"syntology":{"n":17,"n_ran":7,"n_unverified":10,"n_pointer_only":0}},{"url":"/paper/vid-tldr-training-free-token-merging-for","title":"vid-TLDR: Training Free Token merging for Light-weight Video Transformer","date":"2024-03-20","arxiv_id":"2403.13347","repositories_listed":1,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/howtocaption-prompting-llms-to-transform","title":"HowToCaption: Prompting LLMs to Transform Video Annotations at Scale","date":"2023-10-07","arxiv_id":"2310.04900","repositories_listed":1,"syntology":{"n":5,"n_ran":4,"n_unverified":1,"n_pointer_only":5}},{"url":"/paper/one-for-all-video-conversation-is-feasible","title":"BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning","date":"2023-09-27","arxiv_id":"2309.15785","repositories_listed":1,"syntology":{"n":4,"n_ran":1,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/unmasked-teacher-towards-training-efficient","title":"Unmasked Teacher: Towards Training-Efficient Video Foundation Models","date":"2023-03-28","arxiv_id":"2303.16058","repositories_listed":1,"syntology":{"n":8,"n_ran":3,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/seeing-what-you-miss-vision-language-pre","title":"Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning","date":"2022-11-24","arxiv_id":"2211.13437","repositories_listed":1,"syntology":null},{"url":"/paper/clover-towards-a-unified-video-language","title":"Clover: Towards A Unified Video-Language Alignment and Fusion Model","date":"2022-07-16","arxiv_id":"2207.07885","repositories_listed":1,"syntology":null},{"url":"/paper/miles-visual-bert-pre-training-with-injected","title":"MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval","date":"2022-04-26","arxiv_id":"2204.12408","repositories_listed":1,"syntology":null},{"url":"/paper/socratic-models-composing-zero-shot","title":"Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language","date":"2022-04-01","arxiv_id":"2204.00598","repositories_listed":1,"syntology":null},{"url":"/paper/everything-at-once-multi-modal-fusion-1","title":"Everything at Once - Multi-Modal Fusion Transformer for Video Retrieval","date":"2022-01-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/align-and-prompt-video-and-language-pre","title":"Align and Prompt: Video-and-Language Pre-training with Entity Prompts","date":"2021-12-17","arxiv_id":"2112.09583","repositories_listed":1,"syntology":null},{"url":"/paper/everything-at-once-multi-modal-fusion","title":"Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval","date":"2021-12-08","arxiv_id":"2112.04446","repositories_listed":1,"syntology":null},{"url":"/paper/object-aware-video-language-pre-training-for","title":"Object-aware Video-language Pre-training for Retrieval","date":"2021-12-01","arxiv_id":"2112.00656","repositories_listed":1,"syntology":null},{"url":"/paper/violet-end-to-end-video-language-transformers","title":"VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling","date":"2021-11-24","arxiv_id":"2111.12681","repositories_listed":1,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":2}}],"syntology_records":19,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}