| COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training |
1 |
2 |
2 Dec 2024 |
not harvested |
| 3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting |
1 |
1 |
26 Apr 2024 |
ran 10 of 10 samples (0 unverified) |
| Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning |
1 |
1 |
16 Apr 2024 |
not harvested |
| M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining |
1 |
1 |
29 Jan 2024 |
not harvested |
| InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks |
2 |
4 |
21 Dec 2023 |
ran 2 of 2 samples (0 unverified; 2 pointer-only for licence) |
| Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis |
1 |
1 |
21 Sep 2023 |
not harvested |
| VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset |
2 |
2 |
29 May 2023 |
ran 15 of 42 samples (27 unverified) |
| ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities |
2 |
1 |
18 May 2023 |
ran 2 of 7 samples (5 unverified) |
| Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers |
2 |
1 |
11 May 2023 |
ran 5 of 8 samples (3 unverified) |
| MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks |
1 |
1 |
29 Mar 2023 |
ran 3 of 3 samples (0 unverified) |
| Plug-and-Play Regulators for Image-Text Matching |
1 |
2 |
23 Mar 2023 |
not harvested |
| BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
17 |
4 |
30 Jan 2023 |
ran 4 of 8 samples (4 unverified; 1 pointer-only for licence) |
| HADA: A Graph-based Amalgamation Framework in Image-text Retrieval |
2 |
5 |
11 Jan 2023 |
not harvested |
| NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal Embeddings |
1 |
1 |
7 Jan 2023 |
not harvested |
| Position-guided Text Prompt for Vision-Language Pre-training |
1 |
1 |
19 Dec 2022 |
ran 3 of 5 samples (2 unverified) |
| Reproducible scaling laws for contrastive language-image learning |
5 |
1 |
14 Dec 2022 |
ran 3 of 3 samples (0 unverified; 3 pointer-only for licence) |
| X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks |
2 |
2 |
22 Nov 2022 |
ran 2 of 6 samples (4 unverified; 6 pointer-only for licence) |
| AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities |
2 |
1 |
12 Nov 2022 |
ran 2 of 11 samples (9 unverified) |
| Text-Only Training for Image Captioning using Noise-Injected CLIP |
4 |
1 |
1 Nov 2022 |
ran 4 of 5 samples (1 unverified; 1 pointer-only for licence) |
| Dissecting Deep Metric Learning Losses for Image-Text Retrieval |
2 |
1 |
21 Oct 2022 |
not harvested |
| A Comprehensive Study on Large-Scale Graph Training: Benchmarking and Rethinking |
2 |
1 |
14 Oct 2022 |
ran 4 of 6 samples (2 unverified) |
| ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training |
1 |
3 |
30 Sep 2022 |
not harvested |
| OmniVL:One Foundation Model for Image-Language and Video-Language Tasks |
0 |
1 |
15 Sep 2022 |
not harvested |
| Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks |
2 |
2 |
22 Aug 2022 |
not harvested |
| Language Models are General-Purpose Interfaces |
1 |
1 |
13 Jun 2022 |
not harvested |
| CoCa: Contrastive Captioners are Image-Text Foundation Models |
6 |
1 |
4 May 2022 |
ran 9 of 17 samples (8 unverified) |
| Flamingo: a Visual Language Model for Few-Shot Learning |
5 |
1 |
29 Apr 2022 |
ran 18 of 24 samples (6 unverified; 7 pointer-only for licence) |
| ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval |
0 |
1 |
31 Mar 2022 |
not harvested |
| Florence: A New Foundation Model for Computer Vision |
2 |
1 |
22 Nov 2021 |
not harvested |
| Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts |
1 |
2 |
16 Nov 2021 |
ran 1 of 1 samples (0 unverified) |