{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/seeing-out-of-the-box-end-to-end-pre-training","title":"Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning","arxiv_id":"2104.03135","date":"2021-04-07","proceeding":"CVPR 2021 1","authors":["Zhicheng Huang","Zhaoyang Zeng","Yupan Huang","Bei Liu","Dongmei Fu","Jianlong Fu"],"abstract":"We study joint learning of Convolutional Neural Network (CNN) and Transformer for vision-language pre-training (VLPT) which aims to learn cross-modal alignments from millions of image-text pairs. State-of-the-art approaches extract salient image regions and align regions with words step-by-step. As region-based visual features usually represent parts of an image, it is challenging for existing vision-language models to fully understand the semantics from paired natural languages. In this paper, we propose SOHO to \"See Out of tHe bOx\" that takes a whole image as input, and learns vision-language representation in an end-to-end manner. SOHO does not require bounding box annotations which enables inference 10 times faster than region-based approaches. In particular, SOHO learns to extract comprehensive yet compact image features through a visual dictionary (VD) that facilitates cross-modal understanding. VD is designed to represent consistent visual abstractions of similar semantics. It is updated on-the-fly and utilized in our proposed pre-training task Masked Visual Modeling (MVM). We conduct experiments on four well-established vision-language tasks by following standard VLPT settings. In particular, SOHO achieves absolute gains of 2.0% R@1 score on MSCOCO text retrieval 5k test split, 1.5% accuracy on NLVR$^2$ test-P split, 6.7% accuracy on SNLI-VE test split, respectively.","url_abs":"https://arxiv.org/abs/2104.03135v2","url_pdf":"https://arxiv.org/pdf/2104.03135v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"seeing-out-of-the-box-end-to-end-pre-training","repo_url":"https://github.com/researchmm/soho","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"seeing-out-of-the-box-end-to-end-pre-training","repo_url":"https://github.com/PasserBy4/mypretraining","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"seeing-out-of-the-box-end-to-end-pre-training","repo_url":"https://github.com/PasserBy4/pretraining","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"},{"task_slug":"visual-entailment","task_name":"Visual Entailment"},{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"soho","method_name":"SOHO"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-entailment-on-snli-ve-test","task":"Visual Entailment","dataset":"SNLI-VE test","model":"SOHO","rank_in_archive_order":5,"of":8,"metrics":{"Accuracy":"84.95"},"uses_additional_data":false},{"leaderboard":"/sota/visual-entailment-on-snli-ve-val","task":"Visual Entailment","dataset":"SNLI-VE val","model":"SOHO","rank_in_archive_order":5,"of":9,"metrics":{"Accuracy":"85.00"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-nlvr2-dev","task":"Visual Reasoning","dataset":"NLVR2 Dev","model":"SOHO","rank_in_archive_order":12,"of":15,"metrics":{"Accuracy":"76.37"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-nlvr2-test","task":"Visual Reasoning","dataset":"NLVR2 Test","model":"SOHO","rank_in_archive_order":12,"of":14,"metrics":{"Accuracy":"77.32"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2104.03135","atlas_url":"https://app.syntology.ai/?focus=2104.03135","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}