Methods › Computer Vision › Vision and Language Pre-Trained Models › FLAVA

FLAVA

11 papers tagged archive 2025-07-28

Introduced by Amanpreet Singh et al. in FLAVA: A Foundational Language And Vision Alignment Model

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

FLAVA aims at building a single holistic universal model that targets all modalities at once. FLAVA is a language vision alignment model that learns strong representations from multimodal data (image-text pairs) and unimodal data (unpaired images and text). The model consists of an image encode transformer to capture unimodal image representations, a text encoder transformer to process unimodal text information, and a multimodal encode transformer that takes as input the encoded unimodal image and text and integrates their representations for multimodal reasoning. During pretraining, masked image modeling (MIM) and mask language modeling (MLM) losses are applied onto the image and text encoders over a single image or a text piece, respectively, while contrastive, masked multimodal modeling (MMM), and image-text matching (ITM) loss are used over paired image-text data. For downstream tasks, classification heads are applied on the outputs from the image, text, and multimodal encoders respectively for visual recognition, language understanding, and multimodal reasoning tasks It can be applied to broad scope of tasks from three domains (visual recognition, language understanding, and multimodal reasoning) under a common transformer model architecture.

PaperSource

Papers archive 2025-07-28

11 shown of 11, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 25 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modelling4
Language Modeling3
Attribute1
Continual Learning1
Decision Making1
Image Captioning1
Image Retrieval1
Image-text Retrieval1
Image-to-Text Retrieval1
Language Model Evaluation1
Large Language Model1
Object1
Optical Character Recognition1
Optical Character Recognition (OCR)1
Representation Learning1
Retrieval1
Text Retrieval1
Unity1
Video Alignment1
Visual Reasoning1

Usage over time archive 2025-07-28

Papers per year tagged with FLAVA: 2021 to 2024, peak 5 5 0 2021: 1 paper 2021 2022: 1 paper 2022 2023: 4 papers 2023 2024: 5 papers 2024
Papers per year the archive tags with this method, by the paper's archive date (11 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections