Methods › Computer Vision › Vision and Language Pre-Trained Models › OSCAR
OSCAR
Introduced by Xiujun Li et al. in Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
OSCAR is a new learning method that uses object tags detected in images as anchor points to ease the learning of image-text alignment. The model take a triple as input (word-tag-region) and pre-trained with two losses (masked token loss over words and tags, and a contrastive loss between tags and others). OSCAR represents an image-text pair into semantic space via dictionary lookup. Object tags are used as anchor points to align image regions with word embeddings of pre-trained language models. The model is then fine-tuned for understanding and generation tasks.
Papers archive 2025-07-28
30 shown of 36, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
OSCAR: One-Step Diffusion Codec for Image Compression Across Multiple Bit-rates 22 May 2025 · 1 repository · arXiv:2505.16091
-
OSCAR: Online Soft Compression And Reranking 17 Mar 2025 · 0 repositories · arXiv:2504.07109
-
OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking 7 Mar 2025 · 0 repositories · arXiv:2503.05962
-
One-Shot Federated Learning with Classifier-Free Diffusion Models 12 Feb 2025 · 0 repositories · arXiv:2502.08488
-
Longitudinal Abuse and Sentiment Analysis of Hollywood Movie Dialogues using LLMs 20 Jan 2025 · 0 repositories · arXiv:2501.13948
-
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training 28 Oct 2024 · 0 repositories · arXiv:2410.20796
-
OSCAR: Operating System Control via State-Aware Reasoning and Re-Planning 24 Oct 2024 · 0 repositories · arXiv:2410.18963
-
Towards Fine-Grained Webpage Fingerprinting at Scale 6 Sep 2024 · 0 repositories · arXiv:2409.04341
-
IKUN for WMT24 General MT Task: LLMs Are here for Multilingual Machine Translation 21 Aug 2024 · 0 repositories · arXiv:2408.11512
-
Tropical Expressivity of Neural Networks 30 May 2024 · 1 repository · arXiv:2405.20174
-
Strong Screening Rules for Group-based SLOPE Models 24 May 2024 · 1 repository · arXiv:2405.15357
-
Building a Large Japanese Web Corpus for Large Language Models 27 Apr 2024 · 0 repositories · arXiv:2404.17733
-
Automated Model Selection for Generalized Linear Models 25 Apr 2024 · 0 repositories · arXiv:2404.16560
-
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages 13 Mar 2024 · 0 repositories · arXiv:2403.08693
-
OSCaR: Object State Captioning and State Change Representation 27 Feb 2024 · 1 repository · arXiv:2402.17128
-
GlotScript: A Resource and Tool for Low Resource Writing System Identification 23 Sep 2023 · 1 repository · arXiv:2309.13320Syntology ran 0 of 2 samples · 2 unverified
-
A Unified Framework for Pattern Recovery in Penalized and Thresholded Estimation and its Geometry 19 Jul 2023 · 0 repositories · arXiv:2307.10158
-
One-Stage Cascade Refinement Networks for Infrared Small Target Detection 16 Dec 2022 · 2 repositories · arXiv:2212.08472
-
RobBERT-2022: Updating a Dutch Language Model to Account for Evolving Language Use 15 Nov 2022 · 0 repositories · arXiv:2211.08192
-
VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations 1 Jul 2022 · 1 repository · arXiv:2207.00221
-
Impact of Tokenization on Language Models: An Analysis for Turkish 19 Apr 2022 · 0 repositories · arXiv:2204.08832
-
WuDaoMM: A large-scale Multi-Modal Dataset for Pre-training models 22 Mar 2022 · 0 repositories · arXiv:2203.11480
-
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus 17 Jan 2022 · 0 repositories · arXiv:2201.06642
-
Impact of Tokenization on Language Models: An Analysis for Turkish 16 Nov 2021 · 0 repositories
-
PAGnol: An Extra-Large French Generative Model 16 Oct 2021 · 0 repositories · arXiv:2110.08554
-
OSCAR: Data-Driven Operational Space Control for Adaptive and Robust Robot Manipulation 2 Oct 2021 · 1 repository · arXiv:2110.00704
-
Transliteration: A Simple Technique For Improving Multilingual Language Modeling 29 Sep 2021 · 0 repositories
-
A Thorough Review on Recent Deep Learning Methodologies for Image Captioning 28 Jul 2021 · 0 repositories · arXiv:2107.13114
-
Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic Testing 19 Jun 2021 · 1 repository
-
How could Neural Networks understand Programs? 10 May 2021 · 1 repository · arXiv:2105.04297
Tasks archive 2025-07-28
20 shown of 55 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections