Methods › Computer Vision › Vision and Language Pre-Trained Models › SOHO
SOHO
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
SOHO (“See Out of tHe bOx”) that takes a whole image as input, and learns vision-language representation in an end-to-end manner. SOHO does not require bounding box annotations which enables inference 10 times faster than region-based approaches. Text embeddings are used to extract textual embedding features. A trainable CNN is used to extract visual representations. SOHO learns to extract comprehensive yet compact image features through a visual dictionary (VD) that facilitates cross-modal understanding. VD is designed to represent consistent visual abstractions of similar semantics. It is updated on-the-fly and utilized in the proposed pre-training task Masked Visual Modeling (MVM).
Papers archive 2025-07-28
5 shown of 5, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
ThermoONet -- a deep learning-based small body thermophysical network: applications to modelling water activity of comets 20 May 2025 · 0 repositories · arXiv:2505.14016
-
Prediction of Geoeffective CMEs Using SOHO Images and Deep Learning 2 Jan 2025 · 0 repositories · arXiv:2501.01011
-
An Ontology for the Social Determinants of Health Domain 15 Nov 2022 · 0 repositories · arXiv:2211.07837
-
A Machine-Learning-Ready Dataset Prepared from the Solar and Heliospheric Observatory Mission 4 Aug 2021 · 0 repositories · arXiv:2108.06394
-
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning 7 Apr 2021 · 3 repositories · arXiv:2104.03135
Tasks archive 2025-07-28
9 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| BIG-bench Machine Learning | 1 |
| Deep Learning | 1 |
| Representation Learning | 1 |
| Retrieval | 1 |
| Text Retrieval | 1 |
| Transfer Learning | 1 |
| Visual Entailment | 1 |
| Visual Reasoning | 1 |
| global-optimization | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections