{"url":"/dataset/ego4d","name":"Ego4D","full_name":null,"description_markdown":"Ego4D is a massive-scale egocentric video dataset and benchmark suite. It offers 3,025 hours of daily life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 855 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, a host of new benchmark challenges are presented, centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, the aim is to push the frontier of first-person perception.\r\n\r\nDescription from: [Facebook AI](https://ai.facebook.com/research/publications/ego4d-unscripted-first-person-video-from-around-the-world-and-a-benchmark-suite-for-egocentric-perception)\r\n\r\nPaper: [Ego4D: Around the World in 3,000 Hours of Egocentric Video](https://ai.facebook.com/research/publications/ego4d-unscripted-first-person-video-from-around-the-world-and-a-benchmark-suite-for-egocentric-perception)\r\n\r\nGitHub: [https://github.com/EGO4D](https://github.com/EGO4D)","description_withheld":null,"homepage":"https://ego4d-data.org/","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":{"name":"Custom","url":"https://ego4d-data.org/#download"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[{"name":"Temporal Action Localization","url":"/task/action-recognition","datasets_with_task":"/datasets/task/action-recognition"},{"name":"Action Anticipation","url":"/task/action-anticipation","datasets_with_task":"/datasets/task/action-anticipation"},{"name":"Natural Language Queries","url":"/task/natural-language-queries","datasets_with_task":"/datasets/task/natural-language-queries"},{"name":"Moment Queries","url":"/task/moment-queries","datasets_with_task":"/datasets/task/moment-queries"},{"name":"Object State Change Classification","url":"/task/object-state-change-classification","datasets_with_task":"/datasets/task/object-state-change-classification"},{"name":"State Change Object Detection","url":"/task/state-change-object-detection","datasets_with_task":"/datasets/task/state-change-object-detection"},{"name":"Short-term Object Interaction Anticipation","url":"/task/short-term-object-interaction-anticipation","datasets_with_task":"/datasets/task/short-term-object-interaction-anticipation"},{"name":"Future Hand Prediction","url":"/task/future-hand-prediction","datasets_with_task":"/datasets/task/future-hand-prediction"},{"name":"Long Term Action Anticipation","url":"/task/long-term-action-anticipation","datasets_with_task":"/datasets/task/long-term-action-anticipation"}],"languages":[],"variants":["Ego4D","Ego4D MQ val","Ego4D MQ test"],"data_loaders":[],"num_papers_in_archive":32,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/natural-language-queries-on-ego4d","task":"Natural Language Queries","dataset_variant":"Ego4D","rows":10,"metrics":["R@1 Mean(0.3 and 0.5)","R@1 IoU=0.3","R@1 IoU=0.5","R@5 IoU=0.3","R@5 IoU=0.5"],"first_row_in_archive_order":{"model":"EgoVideo","paper":"/paper/egovideo-exploring-egocentric-foundation","metrics":{"R@1 IoU=0.3":"28.05","R@1 IoU=0.5":"19.31","R@1 Mean(0.3 and 0.5)":"23.68","R@5 IoU=0.3":"44.16","R@5 IoU=0.5":"31.37"},"code_links":[{"title":"opengvlab/egovideo","url":"https://github.com/opengvlab/egovideo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/short-term-object-interaction-anticipation-on","task":"Short-term Object Interaction Anticipation","dataset_variant":"Ego4D","rows":4,"metrics":["Overall (Top5 mAP)","Noun (Top5 mAP)","Noun+Verb(Top5 mAP)","Noun+TTC (Top5 mAP)"],"first_row_in_archive_order":{"model":"EgoVideo","paper":"/paper/egovideo-exploring-egocentric-foundation","metrics":{"Noun (Top5 mAP)":"31.08","Noun+TTC (Top5 mAP)":"12.41","Noun+Verb(Top5 mAP)":"16.18","Overall (Top5 mAP)":"7.21"},"code_links":[{"title":"opengvlab/egovideo","url":"https://github.com/opengvlab/egovideo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/future-hand-prediction-on-ego4d","task":"Future Hand Prediction","dataset_variant":"Ego4D","rows":1,"metrics":["Disp(Total)","M.Disp(Left)","C.Disp(Left)","M.Disp(Right)","C.Disp(Right)"],"first_row_in_archive_order":{"model":"InternVideo","paper":"/paper/internvideo-ego4d-a-pack-of-champion","metrics":{"C.Disp(Left)":"53.33","C.Disp(Right)":"53.37","Disp(Total)":"196.8","M.Disp(Left)":"43.25","M.Disp(Right)":"46.25"},"code_links":[{"title":"opengvlab/ego4d-eccv2022-solutions","url":"https://github.com/opengvlab/ego4d-eccv2022-solutions"},{"title":"jonnys1226/ego4d_asl","url":"https://github.com/jonnys1226/ego4d_asl"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/state-change-object-detection-on-ego4d","task":"State Change Object Detection","dataset_variant":"Ego4D","rows":1,"metrics":["AP","AP50","AP75"],"first_row_in_archive_order":{"model":"InternVideo","paper":"/paper/internvideo-ego4d-a-pack-of-champion","metrics":{"AP":"37.19","AP50":"55.97","AP75":"38.44"},"code_links":[{"title":"opengvlab/ego4d-eccv2022-solutions","url":"https://github.com/opengvlab/ego4d-eccv2022-solutions"},{"title":"jonnys1226/ego4d_asl","url":"https://github.com/jonnys1226/ego4d_asl"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/temporal-action-localization-on-ego4d-mq-test","task":"Temporal Action Localization","dataset_variant":"Ego4D MQ test","rows":1,"metrics":["Average mAP","Recall@1x (tIoU=0.5)"],"first_row_in_archive_order":{"model":"ActionFormer (SlowFast+Omnivore+EgoVLP)","paper":"/paper/where-a-strong-backbone-meets-strong-features","metrics":{"Average mAP":"21.76","Recall@1x (tIoU=0.5)":"42.54"},"code_links":[{"title":"happyharrycn/actionformer_release","url":"https://github.com/happyharrycn/actionformer_release"},{"title":"showlab/egovlp","url":"https://github.com/showlab/egovlp"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/temporal-action-localization-on-ego4d-mq-val","task":"Temporal Action Localization","dataset_variant":"Ego4D MQ val","rows":1,"metrics":["Average mAP","Recall@1x (tIoU=0.5)"],"first_row_in_archive_order":{"model":"ActionFormer (SlowFast+Omnivore+EgoVLP)","paper":"/paper/where-a-strong-backbone-meets-strong-features","metrics":{"Average mAP":"21.4","Recall@1x (tIoU=0.5)":"38.73"},"code_links":[{"title":"happyharrycn/actionformer_release","url":"https://github.com/happyharrycn/actionformer_release"},{"title":"showlab/egovlp","url":"https://github.com/showlab/egovlp"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/decafnet-delegate-and-conquer-for-efficient","title":"DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos","date":"2025-05-22","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/short-term-object-interaction-anticipation","title":"Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge","date":"2024-07-08","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/egovideo-exploring-egocentric-foundation","title":"EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation","date":"2024-06-26","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":15,"samples_ran":11,"samples_unverified":4,"pointer_only_for_licence":15,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unimd-towards-unifying-moment-retrieval-and","title":"UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection","date":"2024-04-07","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/rgnet-a-unified-retrieval-and-grounding","title":"RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos","date":"2023-12-11","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/egovlpv2-egocentric-video-language-pre","title":"EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone","date":"2023-07-11","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/guided-attention-for-next-active-object-ego4d","title":"Guided Attention for Next Active Object @ EGO4D STA Challenge","date":"2023-05-25","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/internvideo-ego4d-a-pack-of-champion","title":"InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges","date":"2022-11-17","rows_on_this_dataset":4,"code_links":2,"syntology":null},{"paper":"/paper/where-a-strong-backbone-meets-strong-features","title":"Where a Strong Backbone Meets Strong Features -- ActionFormer for Ego4D Moment Queries Challenge","date":"2022-11-16","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/reler-zju-alibaba-submission-to-the-ego4d","title":"ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022","date":"2022-07-01","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":19,"samples_ran":6,"samples_unverified":13,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/egocentric-video-language-pretraining","title":"Egocentric Video-Language Pretraining","date":"2022-06-03","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":37,"samples_ran":20,"samples_unverified":17,"pointer_only_for_licence":17,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}