Datasets › NExT-QA › Papers where code ran, page 1

NExT-QA

Papers archive 2025-07-28

papers with a benchmark row: 64 · with a code link: 52 · where Syntology ran a sample: 35 (30 with a run with no instrument failure, 5 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (35 of 64 with a benchmark row: 30 with a run with no instrument failure, 5 where every run was a failure of Syntology's instrument)

Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.

Page 1 of 1: papers 1 to 35 of the 35 papers with a benchmark row here where Syntology ran at least one harvested sample (30 with a run with no instrument failure, 5 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 174. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
Agentic Keyframe Search for Video Question Answering 1 1 20 Mar 2025 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 1 1 22 Jan 2025 official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (2 pointer-only for licence)
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 official (archive's flag): 7 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (12 pointer-only for licence)
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 1 1 6 Dec 2024 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence)
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 1 1 19 Sep 2024 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 8 1 18 Sep 2024 official (archive's flag): 4 ran · 12 ran (of which 1 constructed an object rather than computing a result; 10 with no instrument failure: 4 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (6 pointer-only for licence)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 3 3 10 Jul 2024 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (4 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 1 30 Jun 2024 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
Long Context Transfer from Language to Vision 2 1 24 Jun 2024 official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (5 pointer-only for licence)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA 1 1 13 Jun 2024 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 3 1 11 Jun 2024 community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 5 unverified (8 pointer-only for licence)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos 1 1 29 May 2024 official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 2 27 Mar 2024 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified
Understanding Long Videos with Multimodal Language Models 1 1 25 Mar 2024 official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
Language Repository for Long Video Understanding 1 1 21 Mar 2024 official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge 2 1 25 Feb 2024 official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion 1 1 8 Feb 2024 official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering 1 1 3 Jan 2024 official (archive's flag): 14 ran · 14 ran (of which 2 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 1 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (7 pointer-only for licence)
A Simple LLM Framework for Long-Range Video Question-Answering 1 2 28 Dec 2023 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
ViLA: Efficient Video-Language Alignment for Video Question Answering 1 2 13 Dec 2023 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 4 28 Nov 2023 official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (2 pointer-only for licence)
Vamos: Versatile Action Models for Video Understanding 1 1 22 Nov 2023 official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (5 pointer-only for licence)
PaLI-3 Vision Language Models: Smaller, Faster, Stronger 1 1 13 Oct 2023 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
Mistral 7B 6 1 10 Oct 2023 official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
PaLI-X: On Scaling up a Multilingual Vision and Language Model 2 1 29 May 2023 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 4 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
Paxion: Patching Action Knowledge in Video-Language Foundation Models 1 1 18 May 2023 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
Self-Chained Image-Language Model for Video Localization and Question Answering 1 2 11 May 2023 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (7 pointer-only for licence)
Verbs in Action: Improving verb understanding in video-language models 1 2 13 Apr 2023 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
Contrastive Video Question Answering via Video Graph Transformer 1 2 27 Feb 2023 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
Video Graph Transformer for Video Question Answering 1 2 12 Jul 2022 official (archive's flag): 9 ran · 9 ran (of which 5 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified
Revisiting the "Video" in Video-Language Understanding 1 1 3 Jun 2022 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
Flamingo: a Visual Language Model for Few-Shot Learning 5 2 29 Apr 2022 18 ran (of which 6 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (8 pointer-only for licence)
Video as Conditional Graph Hierarchy for Multi-Granular Question Answering 1 1 12 Dec 2021 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)