{"url":"/method/ofa","slug":"ofa","name":"OFA","full_name":"OFA","full_name_withheld":false,"description_markdown":"In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at https://github.com/OFA-Sys/OFA.","description_state":"present","introduced_year":null,"introduced_by":{"title":"OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework","paper":"/paper/unifying-architectures-tasks-and-modalities","first_author":"Peng Wang","n_authors":10,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/unifying-architectures-tasks-and-modalities"},"source":{"url":"https://arxiv.org/abs/2202.03052v2","title":"OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":32,"archive_num_papers":32,"papers_newest_first":[{"paper":null,"title":"MARCO: Hardware-Aware Neural Architecture Search for Edge Devices with Multi-Agent Reinforcement Learning and Conformal Prediction Filtering","date":"2025-06-16","arxiv_id":"2506.13755","n_code_links":0,"syntology":null},{"paper":null,"title":"Private MEV Protection RPCs: Benchmark Stud","date":"2025-05-26","arxiv_id":"2505.19708","n_code_links":0,"syntology":null},{"paper":null,"title":"Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation","date":"2025-05-21","arxiv_id":"2505.15098","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning Object Focused Attention","date":"2025-04-10","arxiv_id":"2504.08166","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficient Adaptation For Remote Sensing Visual Grounding","date":"2025-03-29","arxiv_id":"2503.23083","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison","date":"2025-02-20","arxiv_id":"2502.14827","n_code_links":0,"syntology":null},{"paper":null,"title":"Analysis of the Order Flow Auction under Proposer-Builder Separation","date":"2025-02-17","arxiv_id":"2502.12026","n_code_links":0,"syntology":null},{"paper":"/paper/memory-optimized-once-for-all-network","title":"Memory-Optimized Once-For-All Network","date":"2024-09-05","arxiv_id":"2409.05900","n_code_links":2,"syntology":null},{"paper":null,"title":"Enhancing Journalism with AI: A Study of Contextualized Image Captioning for News Articles using LLMs and LMMs","date":"2024-08-08","arxiv_id":"2408.04331","n_code_links":0,"syntology":null},{"paper":null,"title":"Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge","date":"2024-07-05","arxiv_id":"2407.04255","n_code_links":0,"syntology":null},{"paper":null,"title":"Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering","date":"2024-06-03","arxiv_id":"2406.01402","n_code_links":0,"syntology":null},{"paper":null,"title":"The Solution for the CVPR2024 NICE Image Captioning Challenge","date":"2024-04-19","arxiv_id":"2404.12739","n_code_links":0,"syntology":null},{"paper":"/paper/angofa-leveraging-ofa-embedding","title":"ANGOFA: Leveraging OFA Embedding Initialization and Synthetic Data for Angolan Language Model","date":"2024-04-03","arxiv_id":"2404.02534","n_code_links":1,"syntology":null},{"paper":"/paper/ofa-a-framework-of-initializing-unseen","title":"OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining","date":"2023-11-15","arxiv_id":"2311.08849","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":null,"title":"The Solution for the CVPR2023 NICE Image Captioning Challenge","date":"2023-10-10","arxiv_id":"2310.06879","n_code_links":0,"syntology":null},{"paper":null,"title":"Lightweight In-Context Tuning for Multimodal Unified Models","date":"2023-10-08","arxiv_id":"2310.05109","n_code_links":0,"syntology":null},{"paper":"/paper/one-for-all-towards-training-one-graph-model","title":"One for All: Towards Training One Graph Model for All Classification Tasks","date":"2023-09-29","arxiv_id":"2310.00149","n_code_links":1,"syntology":{"ran":5,"of":18,"unverified":13,"pointer_only":0}},{"paper":"/paper/physics-inspired-hybrid-attention-for-sar","title":"Physics Inspired Hybrid Attention for SAR Target Recognition","date":"2023-09-27","arxiv_id":"2309.15697","n_code_links":1,"syntology":null},{"paper":"/paper/alip-adaptive-language-image-pre-training","title":"ALIP: Adaptive Language-Image Pre-training with Synthetic Caption","date":"2023-08-16","arxiv_id":"2308.08428","n_code_links":1,"syntology":{"ran":5,"of":10,"unverified":5,"pointer_only":10}},{"paper":"/paper/table-and-image-generation-for-investigating","title":"Table and Image Generation for Investigating Knowledge of Entities in Pre-trained Vision and Language Models","date":"2023-06-03","arxiv_id":"2306.02115","n_code_links":1,"syntology":{"ran":6,"of":11,"unverified":5,"pointer_only":7}},{"paper":null,"title":"OFAR: A Multimodal Evidence Retrieval Framework for Illegal Live-streaming Identification","date":"2023-04-25","arxiv_id":"2304.12608","n_code_links":0,"syntology":null},{"paper":null,"title":"oBERTa: Improving Sparse Transfer Learning via improved initialization, distillation, and pruning regimes","date":"2023-03-30","arxiv_id":"2303.17612","n_code_links":0,"syntology":null},{"paper":"/paper/ofa-2-a-multi-objective-perspective-for-the","title":"OFA$^2$: A Multi-Objective Perspective for the Once-for-All Neural Architecture Search","date":"2023-03-23","arxiv_id":"2303.13683","n_code_links":1,"syntology":null},{"paper":null,"title":"Enhancing Once-For-All: A Study on Parallel Blocks, Skip Connections and Early Exits","date":"2023-02-03","arxiv_id":"2302.01888","n_code_links":0,"syntology":null},{"paper":"/paper/binaryvqa-a-versatile-test-set-to-evaluate","title":"BinaryVQA: A Versatile Test Set to Evaluate the Out-of-Distribution Generalization of VQA Models","date":"2023-01-28","arxiv_id":"2301.12032","n_code_links":1,"syntology":null},{"paper":"/paper/multiinstruct-improving-multi-modal-zero-shot","title":"MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning","date":"2022-12-21","arxiv_id":"2212.10773","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":0}},{"paper":"/paper/nas-lid-efficient-neural-architecture-search","title":"NAS-LID: Efficient Neural Architecture Search with Local Intrinsic Dimension","date":"2022-11-23","arxiv_id":"2211.12759","n_code_links":1,"syntology":null},{"paper":null,"title":"How good are deep models in understanding the generated images?","date":"2022-08-23","arxiv_id":"2208.10760","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving Predictive Performance and Calibration by Weight Fusion in Semantic Segmentation","date":"2022-07-22","arxiv_id":"2207.11211","n_code_links":0,"syntology":null},{"paper":"/paper/does-interference-exist-when-training-a-once","title":"Does Interference Exist When Training a Once-For-All Network?","date":"2022-04-20","arxiv_id":"2204.09210","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/question-answering","name":"Question Answering","papers":7},{"task":"/task/visual-question-answering-1","name":"Visual Question Answering","papers":7},{"task":"/task/all","name":"All","papers":6},{"task":"/task/image-captioning","name":"Image Captioning","papers":5},{"task":"/task/retrieval","name":"Retrieval","papers":5},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":5},{"task":"/task/language-modelling","name":"Language Modelling","papers":4},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":4},{"task":"/task/visual-grounding","name":"Visual Grounding","papers":4},{"task":"/task/diversity","name":"Diversity","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3},{"task":"/task/object","name":"Object","papers":3},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":3},{"task":"/task/articles","name":"Articles","papers":2},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":2},{"task":"/task/image-generation","name":"Image Generation","papers":2},{"task":"/task/in-context-learning","name":"In-Context Learning","papers":2},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":2},{"task":"/task/visual-entailment","name":"Visual Entailment","papers":2},{"task":"/task/zero-shot-learning","name":"Zero-Shot Learning","papers":2}],"tasks_shown":20,"n_tasks":64,"usage_by_year":[{"year":"2022","papers":7},{"year":"2023","papers":12},{"year":"2024","papers":6},{"year":"2025","papers":7}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/ofa"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}