{"url":"/method/oner","slug":"oner","name":"OneR","full_name":"One Representation","full_name_withheld":false,"description_markdown":"In the OneR method, model input can be one of image, text or image+text, and CMC objective is combined with the traditional image-text contrastive (ITC) loss. Masked modeling is also carried out for all three input types (i.e., image, text and multi-modal). This framework employs no modality-specific architectural component except for the initial token embedding layer, making our model generic and modality-agnostic with minimal inductive bias.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Unifying Vision-Language Representation Space with Single-tower Transformer","paper":"/paper/unifying-vision-language-representation-space","first_author":"Jiho Jang","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/unifying-vision-language-representation-space"},"source":{"url":"https://arxiv.org/abs/2211.11153v1","title":"Unifying Vision-Language Representation Space with Single-tower Transformer","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":null,"title":"ONER: Online Experience Replay for Incremental Anomaly Detection","date":"2024-12-05","arxiv_id":"2412.03907","n_code_links":0,"syntology":null},{"paper":"/paper/unifying-vision-language-representation-space","title":"Unifying Vision-Language Representation Space with Single-tower Transformer","date":"2022-11-21","arxiv_id":"2211.11153","n_code_links":0,"syntology":null}],"papers_shown":2,"tasks":[{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/object-localization","name":"Object Localization","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/retrieval","name":"Retrieval","papers":1},{"task":"/task/visual-reasoning","name":"Visual Reasoning","papers":1}],"tasks_shown":6,"n_tasks":6,"usage_by_year":[{"year":"2022","papers":1},{"year":"2024","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/oner"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}