| LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token |
1 |
1 |
7 Jan 2025 |
official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| LinVT: Empower Your Image-level Large Language Model to Understand Videos |
1 |
1 |
6 Dec 2024 |
official (archive's flag): 7 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (12 pointer-only for licence) |
| TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models |
1 |
1 |
17 Nov 2024 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance |
1 |
1 |
4 Nov 2024 |
official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (3 pointer-only for licence) |
| SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models |
1 |
1 |
22 Jul 2024 |
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (6 pointer-only for licence) |
| Tarsier: Recipes for Training and Evaluating Large Video Description Models |
1 |
1 |
30 Jun 2024 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified |
| VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding |
1 |
1 |
13 Jun 2024 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (8 pointer-only for licence) |
| MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding |
1 |
1 |
8 Apr 2024 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens |
2 |
1 |
4 Apr 2024 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified |
| ST-LLM: Large Language Models Are Effective Temporal Learners |
1 |
1 |
30 Mar 2024 |
official (archive's flag): 6 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 1 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (7 pointer-only for licence) |
| An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM |
1 |
1 |
27 Mar 2024 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified |
| Elysium: Exploring Object-level Perception in Videos via MLLM |
1 |
1 |
25 Mar 2024 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (8 pointer-only for licence) |
| vid-TLDR: Training Free Token merging for Light-weight Video Transformer |
1 |
1 |
20 Mar 2024 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified |
| CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios |
1 |
1 |
7 Mar 2024 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (2 pointer-only for licence) |
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization |
1 |
1 |
5 Feb 2024 |
3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (5 pointer-only for licence) |
| Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos |
1 |
1 |
16 Dec 2023 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (7 pointer-only for licence) |
| MVBench: A Comprehensive Multi-modal Video Understanding Benchmark |
3 |
1 |
28 Nov 2023 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (2 pointer-only for licence) |
| LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models |
2 |
2 |
28 Nov 2023 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| Video-LLaVA: Learning United Visual Representation by Alignment Before Projection |
6 |
1 |
16 Nov 2023 |
community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence) |
| BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning |
1 |
2 |
27 Sep 2023 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models |
1 |
3 |
18 Aug 2023 |
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (8 pointer-only for licence) |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding |
4 |
1 |
5 Jun 2023 |
official (archive's flag): 7 ran · 18 ran (of which 7 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 10 where Syntology's instrument failed) · 7 unverified (9 pointer-only for licence) |
| VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset |
2 |
1 |
29 May 2023 |
official (archive's flag): 12 ran · 35 ran (of which 4 constructed an object rather than computing a result; 29 with no instrument failure: 2 honoured, 1 violated, 26 with no contract checked; 6 where Syntology's instrument failed) · 7 unverified (8 pointer-only for licence) |
| MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks |
1 |
1 |
29 Mar 2023 |
3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| Unmasked Teacher: Towards Training-Efficient Video Foundation Models |
1 |
1 |
28 Mar 2023 |
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning |
4 |
2 |
25 Mar 2023 |
official (archive's flag): 1 ran · 12 ran (of which 7 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified |
| mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video |
4 |
2 |
1 Feb 2023 |
community repositories only · 17 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence) |
| InternVideo: General Video Foundation Models via Generative and Discriminative Learning |
2 |
1 |
6 Dec 2022 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified |
| X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks |
2 |
2 |
22 Nov 2022 |
official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (6 pointer-only for licence) |
| Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations |
4 |
2 |
21 Nov 2022 |
official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified |
| Zero-Shot Video Question Answering via Frozen Bidirectional Language Models |
3 |
3 |
16 Jun 2022 |
official: harvested, nothing ran · 14 ran (of which 9 constructed an object rather than computing a result; 14 with no instrument failure: 1 honoured, 1 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 20 unverified (1 pointer-only for licence) |
| Revealing Single Frame Bias for Video-and-Language Learning |
2 |
2 |
7 Jun 2022 |
community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| Flamingo: a Visual Language Model for Few-Shot Learning |
5 |
3 |
29 Apr 2022 |
18 ran (of which 6 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (8 pointer-only for licence) |
| Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering |
1 |
1 |
8 Apr 2019 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |