Methods › Computer Vision › Vision and Language Pre-Trained Models › OSCAR

OSCAR

36 papers tagged archive 2025-07-28

Introduced by Xiujun Li et al. in Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

OSCAR is a new learning method that uses object tags detected in images as anchor points to ease the learning of image-text alignment. The model take a triple as input (word-tag-region) and pre-trained with two losses (masked token loss over words and tags, and a contrastive loss between tags and others). OSCAR represents an image-text pair into semantic space via dictionary lookup. Object tags are used as anchor points to align image regions with word embeddings of pre-trained language models. The model is then fine-tuned for understanding and generation tasks.

PaperSource

Papers archive 2025-07-28

30 shown of 36, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 55 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modeling5
Language Modelling5
Benchmarking3
Image Captioning3
Visual Question Answering (VQA)3
model3
NER2
Object2
Question Answering2
Transfer Learning2
Word Embeddings2
feature selection2
regression2
Caption Generation1
Change Detection1
Clustering1
Cross-Modal Retrieval1
DNN Testing1
Dataset Generation1
Denoising1

Usage over time archive 2025-07-28

Papers per year tagged with OSCAR: 2020 to 2025, peak 10 10 0 2020: 5 papers 2020 2021: 8 papers 2021 2022: 6 papers 2022 2023: 2 papers 2023 2024: 10 papers 2024 2025: 5 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (36 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections