Datasets › ActivityNet-QA › Papers where code ran, page 1
ActivityNet-QA
Papers archive 2025-07-28
papers with a benchmark row: 42 · with a code link: 39 · where Syntology ran a sample: 26 (24 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument) Syntology
Show:
all papers with a benchmark row only where code ran (26 of 42 with a benchmark row: 24 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.
Page 1 of 1: papers 1 to 26 of the 26 papers with a benchmark row here where Syntology ran at least one harvested sample (24 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
A run counts when it is a run with no instrument failure any run, instrument failures included The first hides, in your browser, the papers where every run was a failure of Syntology's instrument; the second shows every paper on this page.
The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 146. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since 2026-09-28 , the first build that kept a record date for it; when this build read Syntology's graph is in the build record .
Paper Code Results Date Samples run Syntology
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
1
1
7 Jan 2025
official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence )
LinVT: Empower Your Image-level Large Language Model to Understand Videos
1
1
6 Dec 2024
official (archive's flag): 7 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (12 pointer-only for licence )
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
1
1
17 Nov 2024
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence )
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
1
1
4 Nov 2024
official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (3 pointer-only for licence )
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
1
1
22 Jul 2024
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (6 pointer-only for licence )
Tarsier: Recipes for Training and Evaluating Large Video Description Models
1
1
30 Jun 2024
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
1
1
13 Jun 2024
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (8 pointer-only for licence )
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
1
1
8 Apr 2024
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence )
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
2
1
4 Apr 2024
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
ST-LLM: Large Language Models Are Effective Temporal Learners
1
1
30 Mar 2024
official (archive's flag): 6 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 1 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (7 pointer-only for licence )
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
1
1
27 Mar 2024
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified
Elysium: Exploring Object-level Perception in Videos via MLLM
1
1
25 Mar 2024
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (8 pointer-only for licence )
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
1
1
7 Mar 2024
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (2 pointer-only for licence )
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
1
1
5 Feb 2024
3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (5 pointer-only for licence )
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
3
2
28 Nov 2023
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (2 pointer-only for licence )
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
2
4
28 Nov 2023
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence )
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
6
2
16 Nov 2023
community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence )
TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding
1
1
29 Oct 2023
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (5 pointer-only for licence )
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
1
2
27 Sep 2023
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence )
Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models
1
3
18 Aug 2023
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (8 pointer-only for licence )
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
4
1
5 Jun 2023
official (archive's flag): 7 ran · 18 ran (of which 7 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 10 where Syntology's instrument failed) · 7 unverified (9 pointer-only for licence )
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
2
1
29 May 2023
official (archive's flag): 12 ran · 35 ran (of which 4 constructed an object rather than computing a result; 29 with no instrument failure: 2 honoured, 1 violated, 26 with no contract checked; 6 where Syntology's instrument failed) · 7 unverified (8 pointer-only for licence )
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
1
1
28 Mar 2023
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence )
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models
3
3
16 Jun 2022
official: harvested, nothing ran · 14 ran (of which 9 constructed an object rather than computing a result; 14 with no instrument failure: 1 honoured, 1 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 20 unverified (1 pointer-only for licence )
Revealing Single Frame Bias for Video-and-Language Learning
2
2
7 Jun 2022
community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence )
Towards Fast Adaptation of Pretrained Contrastive Models for Multi-channel Video-Language Retrieval
1
1
5 Jun 2022
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified