Methods › Computer Vision › Vision and Language Pre-Trained Models › FLAVA
FLAVA
Introduced by Amanpreet Singh et al. in FLAVA: A Foundational Language And Vision Alignment Model
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
FLAVA aims at building a single holistic universal model that targets all modalities at once. FLAVA is a language vision alignment model that learns strong representations from multimodal data (image-text pairs) and unimodal data (unpaired images and text). The model consists of an image encode transformer to capture unimodal image representations, a text encoder transformer to process unimodal text information, and a multimodal encode transformer that takes as input the encoded unimodal image and text and integrates their representations for multimodal reasoning. During pretraining, masked image modeling (MIM) and mask language modeling (MLM) losses are applied onto the image and text encoders over a single image or a text piece, respectively, while contrastive, masked multimodal modeling (MMM), and image-text matching (ITM) loss are used over paired image-text data. For downstream tasks, classification heads are applied on the outputs from the image, text, and multimodal encoders respectively for visual recognition, language understanding, and multimodal reasoning tasks It can be applied to broad scope of tasks from three domains (visual recognition, language understanding, and multimodal reasoning) under a common transformer model architecture.
Papers archive 2025-07-28
11 shown of 11, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
The age of spiritual machines: Language quietus induces synthetic altered states of consciousness in artificial intelligence 30 Sep 2024 · 0 repositories · arXiv:2410.00257
-
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst 14 Jun 2024 · 0 repositories · arXiv:2406.09779
-
Interpreting the structure of multi-object representations in vision encoders 13 Jun 2024 · 0 repositories · arXiv:2406.09067
-
Acquiring Linguistic Knowledge from Multimodal Input 27 Feb 2024 · 0 repositories · arXiv:2402.17936
-
AiGen-FoodReview: A Multimodal Dataset of Machine-Generated Restaurant Reviews and Images on Social Media 16 Jan 2024 · 0 repositories · arXiv:2401.08825
-
WhisBERT: Multimodal Text-Audio Language Modeling on 100M Words 5 Dec 2023 · 1 repository · arXiv:2312.02931
-
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment 12 Oct 2023 · 0 repositories · arXiv:2310.08204
-
MoMo: A shared encoder Model for text, image and multi-Modal representations 11 Apr 2023 · 0 repositories · arXiv:2304.05523
-
Controlling for Stereotypes in Multimodal Language Model Evaluation 3 Feb 2023 · 0 repositories · arXiv:2302.01582
-
VL-Taboo: An Analysis of Attribute-based Zero-shot Capabilities of Vision-Language Models 12 Sep 2022 · 1 repository · arXiv:2209.06103
-
FLAVA: A Foundational Language And Vision Alignment Model 8 Dec 2021 · 4 repositories · arXiv:2112.04482
Tasks archive 2025-07-28
20 shown of 25 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections