Methods › Computer Vision › Vision and Language Pre-Trained Models › OFA

OFA

32 papers tagged archive 2025-07-28

Introduced by Peng Wang et al. in OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at https://github.com/OFA-Sys/OFA.

PaperSource

Papers archive 2025-07-28

30 shown of 32, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 64 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Question Answering7
Visual Question Answering7
All6
Image Captioning5
Retrieval5
Visual Question Answering (VQA)5
Language Modelling4
Neural Architecture Search4
Visual Grounding4
Diversity3
Language Modeling3
Object3
Transfer Learning3
Articles2
Contrastive Learning2
Image Generation2
In-Context Learning2
Knowledge Distillation2
Visual Entailment2
Zero-Shot Learning2

Usage over time archive 2025-07-28

Papers per year tagged with OFA: 2022 to 2025, peak 12 12 0 2022: 7 papers 2022 2023: 12 papers 2023 2024: 6 papers 2024 2025: 7 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (32 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections