Methods › Computer Vision › Vision and Language Pre-Trained Models › OFA
OFA
Introduced by Peng Wang et al. in OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at https://github.com/OFA-Sys/OFA.
Papers archive 2025-07-28
30 shown of 32, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
MARCO: Hardware-Aware Neural Architecture Search for Edge Devices with Multi-Agent Reinforcement Learning and Conformal Prediction Filtering 16 Jun 2025 · 0 repositories · arXiv:2506.13755
-
Private MEV Protection RPCs: Benchmark Stud 26 May 2025 · 0 repositories · arXiv:2505.19708
-
Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation 21 May 2025 · 0 repositories · arXiv:2505.15098
-
Learning Object Focused Attention 10 Apr 2025 · 0 repositories · arXiv:2504.08166
-
Efficient Adaptation For Remote Sensing Visual Grounding 29 Mar 2025 · 0 repositories · arXiv:2503.23083
-
Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison 20 Feb 2025 · 0 repositories · arXiv:2502.14827
-
Analysis of the Order Flow Auction under Proposer-Builder Separation 17 Feb 2025 · 0 repositories · arXiv:2502.12026
-
Memory-Optimized Once-For-All Network 5 Sep 2024 · 2 repositories · arXiv:2409.05900
-
Enhancing Journalism with AI: A Study of Contextualized Image Captioning for News Articles using LLMs and LMMs 8 Aug 2024 · 0 repositories · arXiv:2408.04331
-
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge 5 Jul 2024 · 0 repositories · arXiv:2407.04255
-
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering 3 Jun 2024 · 0 repositories · arXiv:2406.01402
-
The Solution for the CVPR2024 NICE Image Captioning Challenge 19 Apr 2024 · 0 repositories · arXiv:2404.12739
-
ANGOFA: Leveraging OFA Embedding Initialization and Synthetic Data for Angolan Language Model 3 Apr 2024 · 1 repository · arXiv:2404.02534
-
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining 15 Nov 2023 · 1 repository · arXiv:2311.08849Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
-
The Solution for the CVPR2023 NICE Image Captioning Challenge 10 Oct 2023 · 0 repositories · arXiv:2310.06879
-
Lightweight In-Context Tuning for Multimodal Unified Models 8 Oct 2023 · 0 repositories · arXiv:2310.05109
-
One for All: Towards Training One Graph Model for All Classification Tasks 29 Sep 2023 · 1 repository · arXiv:2310.00149Syntology ran 5 of 18 samples · 13 unverified
-
Physics Inspired Hybrid Attention for SAR Target Recognition 27 Sep 2023 · 1 repository · arXiv:2309.15697
-
ALIP: Adaptive Language-Image Pre-training with Synthetic Caption 16 Aug 2023 · 1 repository · arXiv:2308.08428Syntology ran 5 of 10 samples · 5 unverified · 10 pointer-only (licence)
-
Table and Image Generation for Investigating Knowledge of Entities in Pre-trained Vision and Language Models 3 Jun 2023 · 1 repository · arXiv:2306.02115Syntology ran 6 of 11 samples · 5 unverified · 7 pointer-only (licence)
-
OFAR: A Multimodal Evidence Retrieval Framework for Illegal Live-streaming Identification 25 Apr 2023 · 0 repositories · arXiv:2304.12608
-
oBERTa: Improving Sparse Transfer Learning via improved initialization, distillation, and pruning regimes 30 Mar 2023 · 0 repositories · arXiv:2303.17612
-
OFA²: A Multi-Objective Perspective for the Once-for-All Neural Architecture Search 23 Mar 2023 · 1 repository · arXiv:2303.13683
-
Enhancing Once-For-All: A Study on Parallel Blocks, Skip Connections and Early Exits 3 Feb 2023 · 0 repositories · arXiv:2302.01888
-
BinaryVQA: A Versatile Test Set to Evaluate the Out-of-Distribution Generalization of VQA Models 28 Jan 2023 · 1 repository · arXiv:2301.12032
-
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning 21 Dec 2022 · 1 repository · arXiv:2212.10773Syntology ran 0 of 1 samples · 1 unverified
-
NAS-LID: Efficient Neural Architecture Search with Local Intrinsic Dimension 23 Nov 2022 · 1 repository · arXiv:2211.12759
-
How good are deep models in understanding the generated images? 23 Aug 2022 · 0 repositories · arXiv:2208.10760
-
Improving Predictive Performance and Calibration by Weight Fusion in Semantic Segmentation 22 Jul 2022 · 0 repositories · arXiv:2207.11211
-
Does Interference Exist When Training a Once-For-All Network? 20 Apr 2022 · 1 repository · arXiv:2204.09210
Tasks archive 2025-07-28
20 shown of 64 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections