Datasets › STAR Benchmark

STAR Benchmark (Situated Reasoning)

6 Dec 2021 archive 2025-07-28

How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence. STAR Benchmark is a novel benchmark for Situated Reasoning, which provides 60K challenging situated questions in four types of tasks, 140K situated hypergraphs, symbolic situation programs, and logic-grounded diagnosis for real-world video situations. (Data Download, STAR Leaderboard)

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video Question Answering STAR Benchmark VLAP (4 frames) Average Accuracy 67.1 ViLA: Efficient Video-Language Alignment for Video... xijun-cs/vila 17 Compare
Zero-Shot Video Question Answer STAR Benchmark VideoChat2 Accuracy 59.0 MVBench: A Comprehensive Multi-modal Video Understanding... opengvlab/ask-anything +2 4 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 17. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
VidCtx: Context-aware Video Question Answering with Image Models 1 1 23 Dec 2024 not harvested
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering 1 1 1 Apr 2024 not harvested
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering 1 2 3 Jan 2024 ran 9 of 18 samples (9 unverified)
ViLA: Efficient Video-Language Alignment for Video Question Answering 1 1 13 Dec 2023 ran 2 of 2 samples (0 unverified)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 1 28 Nov 2023 ran 7 of 10 samples (3 unverified)
Large Language Models are Temporal and Causal Reasoners for Video Question Answering 1 1 24 Oct 2023 ran 0 of 4 samples (4 unverified)
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model 1 1 27 Sep 2023 not harvested
Self-Chained Image-Language Model for Video Localization and Question Answering 1 2 11 May 2023 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
Learning Situation Hyper-Graphs for Video Question Answering 1 1 18 Apr 2023 ran 8 of 15 samples (7 unverified; 15 pointer-only for licence)
MIST: Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question Answering 1 1 19 Dec 2022 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 2 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Revisiting the "Video" in Video-Language Understanding 1 1 3 Jun 2022 ran 3 of 5 samples (2 unverified)
Flamingo: a Visual Language Model for Few-Shot Learning 5 5 29 Apr 2022 ran 18 of 24 samples (6 unverified; 7 pointer-only for licence)
All in One: Exploring Unified Video-Language Pre-training 1 1 14 Mar 2022 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Apache License 2.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • STAR Benchmark

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections