| CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts |
1 |
1 |
9 May 2024 |
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks |
2 |
1 |
21 Dec 2023 |
community repositories only · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (2 pointer-only for licence) |
| ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities |
2 |
2 |
18 May 2023 |
official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video |
4 |
1 |
1 Feb 2023 |
community repositories only · 17 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence) |
| BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
17 |
18 |
30 Jan 2023 |
community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (1 pointer-only for licence) |
| X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks |
2 |
4 |
22 Nov 2022 |
official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (6 pointer-only for licence) |
| Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training |
3 |
2 |
17 Oct 2022 |
community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| PaLI: A Jointly-Scaled Multilingual Language-Image Model |
1 |
1 |
14 Sep 2022 |
official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified |
| Prompt Tuning for Generative Multimodal Pretrained Models |
1 |
1 |
4 Aug 2022 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| LaKo: Knowledge-driven Visual Question Answering via Late Knowledge-to-Text Injection |
1 |
1 |
26 Jul 2022 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| CoCa: Contrastive Captioners are Image-Text Foundation Models |
6 |
1 |
4 May 2022 |
10 ran (of which 5 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (2 pointer-only for licence) |
| Flamingo: a Visual Language Model for Few-Shot Learning |
5 |
3 |
29 Apr 2022 |
18 ran (of which 6 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (8 pointer-only for licence) |
| OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework |
4 |
2 |
7 Feb 2022 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts |
1 |
1 |
16 Nov 2021 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified |
| Coarse-to-Fine Reasoning for Visual Question Answering |
2 |
1 |
6 Oct 2021 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 1 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence) |
| SimVLM: Simple Visual Language Model Pretraining with Weak Supervision |
2 |
2 |
24 Aug 2021 |
22 ran (of which 8 constructed an object rather than computing a result; 16 with no instrument failure: 1 honoured, 3 violated, 12 with no contract checked; 6 where Syntology's instrument failed) · 15 unverified (32 pointer-only for licence) |
| Align before Fuse: Vision and Language Representation Learning with Momentum Distillation |
6 |
2 |
16 Jul 2021 |
community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision |
6 |
1 |
5 Feb 2021 |
official: harvested, nothing ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; the one sample that ran constructed an object rather than computing a result (1 pointer-only for licence) |
| VinVL: Revisiting Visual Representations in Vision-Language Models |
7 |
2 |
2 Jan 2021 |
community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| Sparse and Continuous Attention Mechanisms |
2 |
2 |
12 Jun 2020 |
official: harvested, nothing ran · 12 ran (of which 9 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (2 pointer-only for licence) |
| Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks |
4 |
1 |
13 Apr 2020 |
official (archive's flag): 1 ran · 13 ran (of which 6 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (1 pointer-only for licence) |
| Visual Commonsense R-CNN |
1 |
2 |
27 Feb 2020 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| Compact Trilinear Interaction for Visual Question Answering |
1 |
1 |
26 Sep 2019 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| UNITER: UNiversal Image-TExt Representation Learning |
7 |
2 |
25 Sep 2019 |
community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| Unified Vision-Language Pre-Training for Image Captioning and VQA |
3 |
1 |
24 Sep 2019 |
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (14 pointer-only for licence) |
| VL-BERT: Pre-training of Generic Visual-Linguistic Representations |
3 |
3 |
22 Aug 2019 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| LXMERT: Learning Cross-Modality Encoder Representations from Transformers |
9 |
2 |
20 Aug 2019 |
community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (3 pointer-only for licence) |
| VisualBERT: A Simple and Performant Baseline for Vision and Language |
10 |
2 |
9 Aug 2019 |
4 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (6 pointer-only for licence) |
| ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks |
11 |
1 |
6 Aug 2019 |
10 ran (of which 6 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 24 unverified (34 pointer-only for licence) |
| Deep Modular Co-Attention Networks for Visual Question Answering |
7 |
2 |
25 Jun 2019 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection |
1 |
2 |
31 Jan 2019 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| Bilinear Attention Networks |
8 |
2 |
21 May 2018 |
official (archive's flag): 7 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 3 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering |
65 |
1 |
25 Jul 2017 |
official: no sample here; runs from other or unrecorded repositories · 9 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 8 where Syntology's instrument failed) · 0 unverified (6 pointer-only for licence) |