| Agentic Keyframe Search for Video Question Answering |
1 |
1 |
20 Mar 2025 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified |
| VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding |
1 |
1 |
22 Jan 2025 |
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (2 pointer-only for licence) |
| LinVT: Empower Your Image-level Large Language Model to Understand Videos |
1 |
1 |
6 Dec 2024 |
official (archive's flag): 7 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (12 pointer-only for licence) |
| Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling |
1 |
1 |
6 Dec 2024 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified |
| TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models |
1 |
1 |
17 Nov 2024 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution |
1 |
1 |
19 Sep 2024 |
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution |
8 |
1 |
18 Sep 2024 |
official (archive's flag): 4 ran · 12 ran (of which 1 constructed an object rather than computing a result; 10 with no instrument failure: 4 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models |
1 |
1 |
22 Jul 2024 |
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (6 pointer-only for licence) |
| LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models |
3 |
3 |
10 Jul 2024 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (4 pointer-only for licence) |
| Tarsier: Recipes for Training and Evaluating Large Video Description Models |
1 |
1 |
30 Jun 2024 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified |
| Long Context Transfer from Language to Vision |
2 |
1 |
24 Jun 2024 |
official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (5 pointer-only for licence) |
| Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA |
1 |
1 |
13 Jun 2024 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs |
3 |
1 |
11 Jun 2024 |
community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 5 unverified (8 pointer-only for licence) |
| VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos |
1 |
1 |
29 May 2024 |
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM |
1 |
2 |
27 Mar 2024 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified |
| Understanding Long Videos with Multimodal Language Models |
1 |
1 |
25 Mar 2024 |
official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| Language Repository for Long Video Understanding |
1 |
1 |
21 Mar 2024 |
official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified |
| Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge |
2 |
1 |
25 Feb 2024 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion |
1 |
1 |
8 Feb 2024 |
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| Glance and Focus: Memory Prompting for Multi-Event Video Question Answering |
1 |
1 |
3 Jan 2024 |
official (archive's flag): 14 ran · 14 ran (of which 2 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 1 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (7 pointer-only for licence) |
| A Simple LLM Framework for Long-Range Video Question-Answering |
1 |
2 |
28 Dec 2023 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| ViLA: Efficient Video-Language Alignment for Video Question Answering |
1 |
2 |
13 Dec 2023 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| MVBench: A Comprehensive Multi-modal Video Understanding Benchmark |
3 |
4 |
28 Nov 2023 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (2 pointer-only for licence) |
| Vamos: Versatile Action Models for Video Understanding |
1 |
1 |
22 Nov 2023 |
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (5 pointer-only for licence) |
| PaLI-3 Vision Language Models: Smaller, Faster, Stronger |
1 |
1 |
13 Oct 2023 |
3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| Mistral 7B |
6 |
1 |
10 Oct 2023 |
official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| PaLI-X: On Scaling up a Multilingual Vision and Language Model |
2 |
1 |
29 May 2023 |
7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 4 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| Paxion: Patching Action Knowledge in Video-Language Foundation Models |
1 |
1 |
18 May 2023 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| Self-Chained Image-Language Model for Video Localization and Question Answering |
1 |
2 |
11 May 2023 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (7 pointer-only for licence) |
| Verbs in Action: Improving verb understanding in video-language models |
1 |
2 |
13 Apr 2023 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified |
| Contrastive Video Question Answering via Video Graph Transformer |
1 |
2 |
27 Feb 2023 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| Video Graph Transformer for Video Question Answering |
1 |
2 |
12 Jul 2022 |
official (archive's flag): 9 ran · 9 ran (of which 5 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified |
| Revisiting the "Video" in Video-Language Understanding |
1 |
1 |
3 Jun 2022 |
3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified |
| Flamingo: a Visual Language Model for Few-Shot Learning |
5 |
2 |
29 Apr 2022 |
18 ran (of which 6 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (8 pointer-only for licence) |
| Video as Conditional Graph Hierarchy for Multi-Granular Question Answering |
1 |
1 |
12 Dec 2021 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |