{"url":"/method/soho","slug":"soho","name":"SOHO","full_name":"SOHO","full_name_withheld":false,"description_markdown":"SOHO (“See Out of tHe bOx”) that takes a whole image as input, and learns vision-language representation in an end-to-end manner. SOHO does not require bounding box annotations which enables inference 10 times faster than region-based approaches. Text embeddings are used to extract textual embedding features. A trainable CNN is used to extract visual representations. SOHO learns to extract comprehensive yet compact image features through a visual dictionary (VD) that facilitates cross-modal understanding. VD is designed to represent consistent visual abstractions of similar semantics. It is updated on-the-fly and utilized in the proposed pre-training task Masked Visual Modeling (MVM).","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2104.03135v2","title":"Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":5,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"ThermoONet -- a deep learning-based small body thermophysical network: applications to modelling water activity of comets","date":"2025-05-20","arxiv_id":"2505.14016","n_code_links":0,"syntology":null},{"paper":null,"title":"Prediction of Geoeffective CMEs Using SOHO Images and Deep Learning","date":"2025-01-02","arxiv_id":"2501.01011","n_code_links":0,"syntology":null},{"paper":null,"title":"An Ontology for the Social Determinants of Health Domain","date":"2022-11-15","arxiv_id":"2211.07837","n_code_links":0,"syntology":null},{"paper":null,"title":"A Machine-Learning-Ready Dataset Prepared from the Solar and Heliospheric Observatory Mission","date":"2021-08-04","arxiv_id":"2108.06394","n_code_links":0,"syntology":null},{"paper":"/paper/seeing-out-of-the-box-end-to-end-pre-training","title":"Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning","date":"2021-04-07","arxiv_id":"2104.03135","n_code_links":3,"syntology":null}],"papers_shown":5,"tasks":[{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":1},{"task":"/task/deep-learning","name":"Deep Learning","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/retrieval","name":"Retrieval","papers":1},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":1},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":1},{"task":"/task/visual-entailment","name":"Visual Entailment","papers":1},{"task":"/task/visual-reasoning","name":"Visual Reasoning","papers":1},{"task":"/task/global-optimization","name":"global-optimization","papers":1}],"tasks_shown":9,"n_tasks":9,"usage_by_year":[{"year":"2021","papers":2},{"year":"2022","papers":1},{"year":"2025","papers":2}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/soho"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}