Datasets › Perception Test
Perception Test
Perception Test is a benchmark designed to evaluate the perception and reasoning skills of multimodal models. It introduces real-world videos designed to show perceptually interesting situations and defines multiple tasks that require understanding of memory, abstract patterns, physics, and semantics – across visual, audio, and text modalities. The benchmark consists of 11.6k videos, 23s average length, filmed by around 100 participants worldwide. The videos are densely annotated with six types of labels: object and point tracks, temporal action and sound segments, multiple-choice video question-answers and grounded video question-answers. The benchmark probes pre-trained models for their transfer capabilities, in a zero-shot / few-shot or fine tuning regime.
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Video Question Answering | Perception Test | Oyrx (34B) Accuracy (Top-1) 71.4 | Oryx MLLM: On-Demand Spatial-Temporal Understanding at... | oryx-mllm/oryx | 6 | Compare |
| Object Tracking | Perception Test | Siam-FC Average IOU 0.66 | Perception Test: A Diagnostic Benchmark for Multimodal... | deepmind/perception_test | 1 | Compare |
| Point Tracking | Perception Test | Static Baseline Average Jaccard 0.36 | Perception Test: A Diagnostic Benchmark for Multimodal... | deepmind/perception_test | 1 | Compare |
Papers archive 2025-07-28
6 shown of 6 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 10. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| BIMBA: Selective-Scan Compression for Long-Range Video Question Answering | 1 | 1 | 12 Mar 2025 | not harvested |
| Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution | 1 | 1 | 19 Sep 2024 | ran 3 of 6 samples (3 unverified) |
| VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs | 3 | 1 | 11 Jun 2024 | ran 7 of 17 samples (10 unverified) |
| TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering | 1 | 1 | 1 Apr 2024 | not harvested |
| InternVideo2: Scaling Foundation Models for Multimodal Video Understanding | 2 | 1 | 22 Mar 2024 | not harvested |
| Perception Test: A Diagnostic Benchmark for Multimodal Video Models | 1 | 3 | 23 May 2023 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- Perception Test
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections