Datasets › EgoSchema

EgoSchema

17 Aug 2023 archive 2025-07-28

EgoSchema is very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human curated multiple choice question answer pairs, spanning over 250 hours of real video data, covering a very broad range of natural human activity and behavior.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 112. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Qwen2.5-Omni Technical Report 1 1 26 Mar 2025 not harvested
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering 1 1 12 Mar 2025 not harvested
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition 1 1 12 Dec 2024 ran 3 of 19 samples (16 unverified)
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension 1 1 20 Nov 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning 1 2 25 Oct 2024 ran 3 of 6 samples (3 unverified)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding 1 1 22 Oct 2024 ran 6 of 11 samples (5 unverified)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 2 30 Jun 2024 ran 2 of 2 samples (0 unverified)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA 1 2 13 Jun 2024 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 3 1 11 Jun 2024 ran 7 of 17 samples (10 unverified)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos 1 2 29 May 2024 ran 11 of 11 samples (0 unverified)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering 1 1 1 Apr 2024 not harvested
Understanding Long Videos with Multimodal Language Models 1 2 25 Mar 2024 ran 4 of 4 samples (0 unverified; 1 pointer-only for licence)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 1 22 Mar 2024 not harvested
Language Repository for Long Video Understanding 1 2 21 Mar 2024 ran 9 of 9 samples (0 unverified)
Video ReCap: Recursive Captioning of Hour-Long Videos 2 1 20 Feb 2024 ran 3 of 3 samples (0 unverified; 1 pointer-only for licence)
A Simple LLM Framework for Long-Range Video Question-Answering 1 4 28 Dec 2023 ran 6 of 6 samples (0 unverified)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding 2 1 4 Dec 2023 ran 7 of 11 samples (4 unverified)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 5 28 Nov 2023 ran 7 of 10 samples (3 unverified)
Vamos: Versatile Action Models for Video Understanding 1 3 22 Nov 2023 ran 7 of 8 samples (1 unverified)
Self-Chained Image-Language Model for Video Localization and Question Answering 1 2 11 May 2023 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality 1 1 27 Apr 2023 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 1 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models 3 1 16 Jun 2022 ran 14 of 34 samples (20 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • EgoSchema
  • EgoSchema (subset)
  • EgoSchema (fullset)

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections