{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/egocentric-video-language-pretraining","title":"Egocentric Video-Language Pretraining","arxiv_id":"2206.01670","date":"2022-06-03","proceeding":null,"authors":["Kevin Qinghong Lin","Alex Jinpeng Wang","Mattia Soldan","Michael Wray","Rui Yan","Eric Zhongcong Xu","Difei Gao","RongCheng Tu","Wenzhe Zhao","Weijie Kong","Chengfei Cai","Hongfa Wang","Dima Damen","Bernard Ghanem","Wei Liu","Mike Zheng Shou"],"abstract":"Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at https://github.com/showlab/EgoVLP.","url_abs":"https://arxiv.org/abs/2206.01670v2","url_pdf":"https://arxiv.org/pdf/2206.01670v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"egocentric-video-language-pretraining","repo_url":"https://github.com/showlab/egovlp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"egocentric-video-language-pretraining","repo_url":"https://github.com/zhaoyue-zephyrus/avion","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"moment-queries","task_name":"Moment Queries"},{"task_slug":"multi-instance-retrieval","task_name":"Multi-Instance Retrieval"},{"task_slug":"natural-language-queries","task_name":"Natural Language Queries"},{"task_slug":"object-state-change-classification","task_name":"Object State Change Classification"},{"task_slug":null,"task_name":"Object State Change Classification on Ego4D"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"temporal-localization","task_name":"Temporal Localization"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"},{"task_slug":"video-summarization","task_name":"Video Summarization"},{"task_slug":"video-text-retrieval","task_name":"Video-Text Retrieval"}],"methods":[{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-on-charades-ego","task":"Action Recognition","dataset":"Charades-Ego","model":"EgoVLP","rank_in_archive_order":4,"of":6,"metrics":{"mAP":"32.1"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-queries-on-ego4d","task":"Natural Language Queries","dataset":"Ego4D","model":"EgoVLP","rank_in_archive_order":9,"of":10,"metrics":{"R@1 IoU=0.3":"10.46","R@1 IoU=0.5":"6.24","R@1 Mean(0.3 and 0.5)":"8.35","R@5 IoU=0.3":"16.76","R@5 IoU=0.5":"11.29"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-egotaskqa","task":"Question Answering","dataset":"EgoTaskQA","model":"EgoVLP","rank_in_archive_order":4,"of":4,"metrics":{"Direct":"42.51"},"uses_additional_data":false},{"leaderboard":"/sota/video-summarization-on-query-focused-video","task":"Video Summarization","dataset":"Query-Focused Video Summarization Dataset","model":"EgoVLP","rank_in_archive_order":2,"of":2,"metrics":{"F1 (avg)":"49.72"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2206.01670","atlas_url":"https://app.syntology.ai/?focus=2206.01670","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2206.01670"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/showlab/EgoVLP","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zhaoyue-zephyrus/avion","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/showlab/egovlp","reach":{"status":"ok"}}],"summary":{"ran":2,"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1},"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"9f03e660263fc07e","entry":"ClipLoss","repo":"zhaoyue-zephyrus/avion","repo_kind":"listed","path":"avion/losses/losses.py","file_url":"https://github.com/zhaoyue-zephyrus/avion/blob/HEAD/avion/losses/losses.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9f03e660263fc07e"}},{"code_sha256_prefix":"13ebd1ac0c5f7ae0","entry":"EgoNCE","repo":"showlab/egovlp","repo_kind":"official","path":"model/loss.py","file_url":"https://github.com/showlab/egovlp/blob/HEAD/model/loss.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"13ebd1ac0c5f7ae0"}},{"code_sha256_prefix":"185c7b880ca6135e","entry":"sim_matrix","repo":"showlab/egovlp","repo_kind":"official","path":"model/model.py","file_url":"https://github.com/showlab/egovlp/blob/HEAD/model/model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"185c7b880ca6135e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}