{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/x-clip-end-to-end-multi-grained-contrastive","title":"X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval","arxiv_id":"2207.07285","date":"2022-07-15","proceeding":null,"authors":["Yiwei Ma","Guohai Xu","Xiaoshuai Sun","Ming Yan","Ji Zhang","Rongrong Ji"],"abstract":"Video-text retrieval has been a crucial and fundamental task in multi-modal research. The development of video-text retrieval has been considerably promoted by large-scale multi-modal contrastive pre-training, which primarily focuses on coarse-grained or fine-grained contrast. However, cross-grained contrast, which is the contrast between coarse-grained representations and fine-grained representations, has rarely been explored in prior research. Compared with fine-grained or coarse-grained contrasts, cross-grained contrast calculate the correlation between coarse-grained features and each fine-grained feature, and is able to filter out the unnecessary fine-grained features guided by the coarse-grained feature during similarity calculation, thus improving the accuracy of retrieval. To this end, this paper presents a novel multi-grained contrastive model, namely X-CLIP, for video-text retrieval. However, another challenge lies in the similarity aggregation problem, which aims to aggregate fine-grained and cross-grained similarity matrices to instance-level similarity. To address this challenge, we propose the Attention Over Similarity Matrix (AOSM) module to make the model focus on the contrast between essential frames and words, thus lowering the impact of unnecessary frames and words on retrieval results. With multi-grained contrast and the proposed AOSM module, X-CLIP achieves outstanding performance on five widely-used video-text retrieval datasets, including MSR-VTT (49.3 R@1), MSVD (50.4 R@1), LSMDC (26.1 R@1), DiDeMo (47.8 R@1) and ActivityNet (46.2 R@1). It outperforms the previous state-of-theart by +6.3%, +6.6%, +11.1%, +6.7%, +3.8% relative improvements on these benchmarks, demonstrating the superiority of multi-grained contrast and AOSM.","url_abs":"https://arxiv.org/abs/2207.07285v2","url_pdf":"https://arxiv.org/pdf/2207.07285v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"x-clip-end-to-end-multi-grained-contrastive","repo_url":"https://github.com/xuguohai/X-CLIP","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"x-clip-end-to-end-multi-grained-contrastive","repo_url":"https://github.com/MindCode-4/code-5/tree/main/x_clip","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"x-clip-end-to-end-multi-grained-contrastive","repo_url":"https://github.com/MindSpore-scientific/code-7/tree/main/X_CLIP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"},{"task_slug":"video-text-retrieval","task_name":"Video-Text Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-retrieval-on-activitynet","task":"Video Retrieval","dataset":"ActivityNet","model":"X-CLIP","rank_in_archive_order":21,"of":31,"metrics":{"text-to-video Mean Rank":"6.8","text-to-video R@1":"46.2","text-to-video R@5":"75.5","video-to-text Mean Rank":"6.4","video-to-text R@1":"46.4","video-to-text R@5":"75.9"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-didemo","task":"Video Retrieval","dataset":"DiDeMo","model":"X-CLIP","rank_in_archive_order":29,"of":40,"metrics":{"text-to-video Mean Rank":"12.6","text-to-video R@1":"47.8","text-to-video R@5":"79.3","video-to-text Mean Rank":"10.5","video-to-text R@1":"47.8","video-to-text R@10":"76.8"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-lsmdc","task":"Video Retrieval","dataset":"LSMDC","model":"X-CLIP","rank_in_archive_order":14,"of":38,"metrics":{"text-to-video R@1":"26.1","video-to-text R@1":"26.9"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-msr-vtt-1ka","task":"Video Retrieval","dataset":"MSR-VTT-1kA","model":"X-CLIP","rank_in_archive_order":20,"of":63,"metrics":{"text-to-video Mean Rank":"12.2","text-to-video Median Rank":"2.0","text-to-video R@1":"49.3","text-to-video R@10":"84.8","text-to-video R@5":"75.8","video-to-text Mean Rank":"8.1","video-to-text Median Rank":"2.0","video-to-text R@1":"48.9","video-to-text R@10":"84.5","video-to-text R@5":"76.8"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-msvd","task":"Video Retrieval","dataset":"MSVD","model":"X-CLIP","rank_in_archive_order":12,"of":24,"metrics":{"text-to-video Mean Rank":"8.4","text-to-video R@1":"50.4","text-to-video R@5":"80.6","video-to-text Mean Rank":"4.2","video-to-text R@1":"66.8","video-to-text R@10":"90.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2207.07285","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2207.07285"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-5/tree/main/x_clip","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/xuguohai/X-CLIP","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific/code-7/tree/main/X_CLIP","reach":null}],"summary":{"ran_fixture":1,"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"a0c23f10479a984e","entry":"init_device","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"a0c23f10479a984e"}},{"code_sha256_prefix":"fe6c2cd80d97afae","entry":"get_args","repo":"xuguohai/X-CLIP","repo_kind":"official","path":"main_xclip.py","file_url":"https://github.com/xuguohai/X-CLIP/blob/HEAD/main_xclip.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fe6c2cd80d97afae"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}