| LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token |
1 |
1 |
7 Jan 2025 |
ran 3 of 3 samples (0 unverified; 1 pointer-only for licence) |
| LinVT: Empower Your Image-level Large Language Model to Understand Videos |
1 |
1 |
6 Dec 2024 |
ran 4 of 12 samples (8 unverified; 12 pointer-only for licence) |
| TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models |
1 |
1 |
17 Nov 2024 |
ran 5 of 11 samples (6 unverified) |
| PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance |
1 |
1 |
4 Nov 2024 |
ran 2 of 9 samples (7 unverified; 3 pointer-only for licence) |
| SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models |
1 |
1 |
22 Jul 2024 |
ran 4 of 6 samples (2 unverified; 6 pointer-only for licence) |
| Tarsier: Recipes for Training and Evaluating Large Video Description Models |
1 |
1 |
30 Jun 2024 |
ran 2 of 2 samples (0 unverified) |
| VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding |
1 |
1 |
13 Jun 2024 |
ran 6 of 8 samples (2 unverified; 8 pointer-only for licence) |
| Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams |
1 |
1 |
12 Jun 2024 |
not harvested |
| PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning |
1 |
1 |
25 Apr 2024 |
ran 0 of 2 samples (2 unverified; 2 pointer-only for licence) |
| Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs |
1 |
2 |
11 Apr 2024 |
not harvested |
| MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding |
1 |
1 |
8 Apr 2024 |
ran 7 of 9 samples (2 unverified) |
| MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens |
2 |
1 |
4 Apr 2024 |
ran 2 of 2 samples (0 unverified) |
| ST-LLM: Large Language Models Are Effective Temporal Learners |
1 |
1 |
30 Mar 2024 |
ran 7 of 11 samples (4 unverified; 3 pointer-only for licence) |
| An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM |
1 |
1 |
27 Mar 2024 |
ran 3 of 3 samples (0 unverified) |
| Elysium: Exploring Object-level Perception in Videos via MLLM |
1 |
1 |
25 Mar 2024 |
ran 7 of 8 samples (1 unverified; 8 pointer-only for licence) |
| CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios |
1 |
1 |
7 Mar 2024 |
ran 6 of 12 samples (6 unverified) |
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization |
1 |
1 |
5 Feb 2024 |
ran 3 of 5 samples (2 unverified; 5 pointer-only for licence) |
| MVBench: A Comprehensive Multi-modal Video Understanding Benchmark |
3 |
2 |
28 Nov 2023 |
ran 7 of 10 samples (3 unverified) |
| LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models |
2 |
4 |
28 Nov 2023 |
ran 3 of 4 samples (1 unverified) |
| Video-LLaVA: Learning United Visual Representation by Alignment Before Projection |
6 |
2 |
16 Nov 2023 |
ran 4 of 7 samples (3 unverified; 1 pointer-only for licence) |
| Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding |
4 |
3 |
14 Nov 2023 |
not harvested |
| Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities |
0 |
1 |
9 Nov 2023 |
not harvested |
| TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding |
1 |
1 |
29 Oct 2023 |
ran 11 of 15 samples (4 unverified) |
| BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning |
1 |
2 |
27 Sep 2023 |
ran 1 of 4 samples (3 unverified) |
| Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models |
1 |
3 |
18 Aug 2023 |
ran 3 of 8 samples (5 unverified; 8 pointer-only for licence) |
| MovieChat: From Dense Token to Sparse Memory for Long Video Understanding |
1 |
2 |
31 Jul 2023 |
not harvested |
| COSA: Concatenated Sample Pretrained Vision-Language Foundation Model |
1 |
1 |
15 Jun 2023 |
not harvested |
| Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models |
2 |
2 |
8 Jun 2023 |
not harvested |
| Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding |
4 |
1 |
5 Jun 2023 |
ran 18 of 25 samples (7 unverified; 9 pointer-only for licence) |
| VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset |
2 |
1 |
29 May 2023 |
ran 15 of 42 samples (27 unverified) |