| 3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting |
1 |
1 |
26 Apr 2024 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks |
2 |
4 |
21 Dec 2023 |
community repositories only · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (2 pointer-only for licence) |
| VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset |
2 |
2 |
29 May 2023 |
official (archive's flag): 12 ran · 35 ran (of which 4 constructed an object rather than computing a result; 29 with no instrument failure: 2 honoured, 1 violated, 26 with no contract checked; 6 where Syntology's instrument failed) · 7 unverified (8 pointer-only for licence) |
| ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities |
2 |
1 |
18 May 2023 |
official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers |
2 |
1 |
11 May 2023 |
official (archive's flag): 5 ran · 5 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified |
| MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks |
1 |
1 |
29 Mar 2023 |
3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
17 |
4 |
30 Jan 2023 |
community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (1 pointer-only for licence) |
| Position-guided Text Prompt for Vision-Language Pre-training |
1 |
1 |
19 Dec 2022 |
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified |
| Reproducible scaling laws for contrastive language-image learning |
5 |
1 |
14 Dec 2022 |
community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks |
2 |
2 |
22 Nov 2022 |
official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (6 pointer-only for licence) |
| AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities |
2 |
1 |
12 Nov 2022 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| Text-Only Training for Image Captioning using Noise-Injected CLIP |
4 |
1 |
1 Nov 2022 |
official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| A Comprehensive Study on Large-Scale Graph Training: Benchmarking and Rethinking |
2 |
1 |
14 Oct 2022 |
official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 4 samples that ran constructed an object rather than computing a result |
| CoCa: Contrastive Captioners are Image-Text Foundation Models |
6 |
1 |
4 May 2022 |
10 ran (of which 5 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (2 pointer-only for licence) |
| Flamingo: a Visual Language Model for Few-Shot Learning |
5 |
1 |
29 Apr 2022 |
18 ran (of which 6 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (8 pointer-only for licence) |
| Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts |
1 |
2 |
16 Nov 2021 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified |
| Align before Fuse: Vision and Language Representation Learning with Momentum Distillation |
6 |
2 |
16 Jul 2021 |
community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| Learning Relation Alignment for Calibrated Cross-modal Retrieval |
1 |
1 |
28 May 2021 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified |
| Learning Transferable Visual Models From Natural Language Supervision |
82 |
1 |
26 Feb 2021 |
official: no sample here; runs from other or unrecorded repositories · 16 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 14 where Syntology's instrument failed) · 4 unverified (16 pointer-only for licence) |
| Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision |
5 |
2 |
11 Feb 2021 |
8 ran (of which 6 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (9 pointer-only for licence) |
| ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision |
6 |
2 |
5 Feb 2021 |
official: harvested, nothing ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; the one sample that ran constructed an object rather than computing a result (1 pointer-only for licence) |
| Unifying Vision-and-Language Tasks via Text Generation |
2 |
1 |
4 Feb 2021 |
official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified; every one of the 2 samples that ran constructed an object rather than computing a result |
| Similarity Reasoning and Filtration for Image-Text Matching |
1 |
2 |
5 Jan 2021 |
official (archive's flag): 10 ran · 10 ran (of which 6 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (12 pointer-only for licence) |
| Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders |
1 |
2 |
12 Aug 2020 |
official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 1 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence) |
| Data Augmentation for Graph Neural Networks |
2 |
1 |
11 Jun 2020 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified |
| IMRAM: Iterative Matching with Recurrent Attention Memory for Cross-Modal Image-Text Retrieval |
1 |
1 |
8 Mar 2020 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| UNITER: UNiversal Image-TExt Representation Learning |
7 |
1 |
25 Sep 2019 |
community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| Unified Vision-Language Pre-Training for Image Captioning and VQA |
3 |
1 |
24 Sep 2019 |
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (14 pointer-only for licence) |
| CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval |
1 |
1 |
12 Sep 2019 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| Visual Semantic Reasoning for Image-Text Matching |
2 |
1 |
6 Sep 2019 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| Stacked Cross Attention for Image-Text Matching |
6 |
2 |
21 Mar 2018 |
official: harvested, nothing ran · 13 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 6 where Syntology's instrument failed) · 3 unverified (1 pointer-only for licence) |
| Graph Attention Networks |
93 |
1 |
30 Oct 2017 |
community repositories only · 61 ran (of which 28 constructed an object rather than computing a result; 52 with no instrument failure: 1 honoured, 2 violated, 49 with no contract checked; 9 where Syntology's instrument failed) · 45 unverified (46 pointer-only for licence) |
| VSE++: Improving Visual-Semantic Embeddings with Hard Negatives |
10 |
1 |
18 Jul 2017 |
official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (4 pointer-only for licence) |
| Inductive Representation Learning on Large Graphs |
20 |
1 |
7 Jun 2017 |
community repositories only · 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (5 pointer-only for licence) |
| Semi-Supervised Classification with Graph Convolutional Networks |
55 |
2 |
9 Sep 2016 |
community repositories only · 39 ran (of which 13 constructed an object rather than computing a result; 34 with no instrument failure: 0 honoured, 1 violated, 33 with no contract checked; 5 where Syntology's instrument failed) · 19 unverified (23 pointer-only for licence) |
| Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models |
2 |
1 |
19 May 2015 |
2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |