| TinyLLaVA: A Framework of Small-scale Large Multimodal Models |
2 |
1 |
22 Feb 2024 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (1 pointer-only for licence) |
| CoLLaVO: Crayon Large Language and Vision mOdel |
1 |
1 |
17 Feb 2024 |
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified |
| SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models |
1 |
1 |
8 Feb 2024 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization |
1 |
1 |
5 Feb 2024 |
3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (5 pointer-only for licence) |
| From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information |
0 |
1 |
31 Jan 2024 |
not harvested |
| MouSi: Poly-Visual-Expert Vision-Language Models |
1 |
1 |
30 Jan 2024 |
not harvested |
| InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model |
1 |
1 |
29 Jan 2024 |
not harvested |
| MoE-LLaVA: Mixture of Experts for Large Vision-Language Models |
3 |
1 |
29 Jan 2024 |
official: harvested, nothing ran · 0 ran · 2 unverified (2 pointer-only for licence) |
| Small Language Model Meets with Reinforced Vision Vocabulary |
0 |
1 |
23 Jan 2024 |
not harvested |
| COCO is "ALL'' You Need for Visual Instruction Fine-tuning |
0 |
1 |
17 Jan 2024 |
not harvested |
| CaMML: Context-Aware Multimodal Learner for Large Models |
1 |
1 |
6 Jan 2024 |
not harvested |
| LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model |
1 |
1 |
4 Jan 2024 |
official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (9 pointer-only for licence) |
| V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs |
1 |
1 |
21 Dec 2023 |
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified |
| Generative Multimodal Models are In-Context Learners |
1 |
1 |
20 Dec 2023 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified |
| Gemini: A Family of Highly Capable Multimodal Models |
1 |
1 |
19 Dec 2023 |
not harvested |
| Silkie: Preference Distillation for Large Visual Language Models |
0 |
1 |
17 Dec 2023 |
not harvested |
| CogAgent: A Visual Language Model for GUI Agents |
3 |
1 |
14 Dec 2023 |
official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (1 pointer-only for licence) |
| Hallucination Augmented Contrastive Learning for Multimodal Large Language Model |
1 |
1 |
12 Dec 2023 |
official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (5 pointer-only for licence) |
| VILA: On Pre-training for Visual Language Models |
3 |
1 |
12 Dec 2023 |
not harvested |
| Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models |
1 |
1 |
11 Dec 2023 |
not harvested |
| OneLLM: One Framework to Align All Modalities with Language |
1 |
1 |
6 Dec 2023 |
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence) |
| Merlin:Empowering Multimodal LLMs with Foresight Minds |
0 |
1 |
30 Nov 2023 |
not harvested |
| ShareGPT4V: Improving Large Multi-Modal Models with Better Captions |
1 |
2 |
21 Nov 2023 |
not harvested |
| Video-LLaVA: Learning United Visual Representation by Alignment Before Projection |
6 |
1 |
16 Nov 2023 |
community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence) |
| SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models |
1 |
1 |
13 Nov 2023 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning |
2 |
1 |
13 Nov 2023 |
not harvested |
| Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision |
1 |
2 |
13 Nov 2023 |
official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (10 pointer-only for licence) |
| InfMLLM: A Unified Framework for Visual-Language Tasks |
2 |
1 |
12 Nov 2023 |
not harvested |
| LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents |
1 |
2 |
9 Nov 2023 |
official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration |
2 |
1 |
7 Nov 2023 |
official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (3 pointer-only for licence) |
| OtterHD: A High-Resolution Multi-modality Model |
1 |
1 |
7 Nov 2023 |
official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| CogVLM: Visual Expert for Pretrained Language Models |
4 |
2 |
6 Nov 2023 |
not harvested |
| Improved Baselines with Visual Instruction Tuning |
9 |
2 |
5 Oct 2023 |
6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (8 pointer-only for licence) |
| DreamLLM: Synergistic Multimodal Comprehension and Creation |
1 |
1 |
20 Sep 2023 |
official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models |
1 |
1 |
18 Sep 2023 |
official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild |
1 |
1 |
14 Sep 2023 |
not harvested |
| Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond |
2 |
2 |
24 Aug 2023 |
official: harvested for another paper · 0 ran · 2 unverified (2 pointer-only for licence) |
| StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data |
1 |
1 |
20 Aug 2023 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models |
2 |
2 |
2 Aug 2023 |
not harvested |
| Emu: Generative Pretraining in Multimodality |
2 |
1 |
11 Jul 2023 |
official: harvested for another paper · 0 ran · 2 unverified |
| Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning |
4 |
1 |
26 Jun 2023 |
not harvested |
| MIMIC-IT: Multi-Modal In-Context Instruction Tuning |
2 |
2 |
8 Jun 2023 |
not harvested |
| LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model |
3 |
1 |
28 Apr 2023 |
official: not harvested · 0 ran · 1 unverified (1 pointer-only for licence) |
| MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models |
6 |
2 |
20 Apr 2023 |
not harvested |
| MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action |
1 |
2 |
20 Mar 2023 |
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| GPT-4 Technical Report |
11 |
5 |
15 Mar 2023 |
community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
17 |
1 |
30 Jan 2023 |
community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (1 pointer-only for licence) |