{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-straightforward-framework-for-video","title":"A Straightforward Framework For Video Retrieval Using CLIP","arxiv_id":"2102.12443","date":"2021-02-24","proceeding":null,"authors":["Jesús Andrés Portillo-Quintero","José Carlos Ortiz-Bayliss","Hugo Terashima-Marín"],"abstract":"Video Retrieval is a challenging task where a text query is matched to a video or vice versa. Most of the existing approaches for addressing such a problem rely on annotations made by the users. Although simple, this approach is not always feasible in practice. In this work, we explore the application of the language-image model, CLIP, to obtain video representations without the need for said annotations. This model was explicitly trained to learn a common space where images and text can be compared. Using various techniques described in this document, we extended its application to videos, obtaining state-of-the-art results on the MSR-VTT and MSVD benchmarks.","url_abs":"https://arxiv.org/abs/2102.12443v2","url_pdf":"https://arxiv.org/pdf/2102.12443v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-straightforward-framework-for-video","repo_url":"https://github.com/Deferf/CLIP_Video_Representation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-retrieval-on-lsmdc","task":"Video Retrieval","dataset":"LSMDC","model":"CLIP","rank_in_archive_order":31,"of":38,"metrics":{"text-to-video Median Rank":"56.5","text-to-video R@1":"11.3","text-to-video R@10":"29.2","text-to-video R@5":"22.7","video-to-text Median Rank":"73","video-to-text R@1":"6.8","video-to-text R@10":"22.1","video-to-text R@5":"16.4"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-msr-vtt","task":"Video Retrieval","dataset":"MSR-VTT","model":"CLIP","rank_in_archive_order":32,"of":40,"metrics":{"text-to-video Median Rank":"10","text-to-video R@1":"21.4","text-to-video R@10":"50.4","text-to-video R@5":"41.1","video-to-text Median Rank":"2","video-to-text R@1":"40.3","video-to-text R@10":"79.2","video-to-text R@5":"69.7"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-msr-vtt-1ka","task":"Video Retrieval","dataset":"MSR-VTT-1kA","model":"CLIP","rank_in_archive_order":47,"of":63,"metrics":{"text-to-video Median Rank":"4","text-to-video R@1":"31.2","text-to-video R@10":"64.2","text-to-video R@5":"53.7","video-to-text Median Rank":"5","video-to-text R@1":"27.2","video-to-text R@10":"62.6","video-to-text R@5":"51.7"},"uses_additional_data":true},{"leaderboard":"/sota/video-retrieval-on-msvd","task":"Video Retrieval","dataset":"MSVD","model":"CLIP","rank_in_archive_order":21,"of":24,"metrics":{"text-to-video Median Rank":"3.0","text-to-video R@1":"37","text-to-video R@10":"73.8","text-to-video R@5":"64.1","video-to-text Median Rank":"1","video-to-text R@1":"59.9","video-to-text R@10":"90.7","video-to-text R@5":"85.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2102.12443","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}