{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ernie-vil-knowledge-enhanced-vision-language","title":"ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph","arxiv_id":"2006.16934","date":"2020-06-30","proceeding":null,"authors":["Fei Yu","Jiji Tang","Weichong Yin","Yu Sun","Hao Tian","Hua Wu","Haifeng Wang"],"abstract":"We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic connections (objects, attributes of objects and relationships between objects) across vision and language, which are essential to vision-language cross-modal tasks. Utilizing scene graphs of visual scenes, ERNIE-ViL constructs Scene Graph Prediction tasks, i.e., Object Prediction, Attribute Prediction and Relationship Prediction tasks in the pre-training phase. Specifically, these prediction tasks are implemented by predicting nodes of different types in the scene graph parsed from the sentence. Thus, ERNIE-ViL can learn the joint representations characterizing the alignments of the detailed semantics across vision and language. After pre-training on large scale image-text aligned datasets, we validate the effectiveness of ERNIE-ViL on 5 cross-modal downstream tasks. ERNIE-ViL achieves state-of-the-art performances on all these tasks and ranks the first place on the VCR leaderboard with an absolute improvement of 3.7%.","url_abs":"https://arxiv.org/abs/2006.16934v3","url_pdf":"https://arxiv.org/pdf/2006.16934v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"referring-expression-comprehension","task_name":"Referring Expression Comprehension"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-question-answering-on-vcr-q-a-test","task":"Visual Question Answering (VQA)","dataset":"VCR (Q-A) test","model":"ERNIE-ViL-large(ensemble of 15 models)","rank_in_archive_order":2,"of":11,"metrics":{"Accuracy":"81.6"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-vcr-q-ar-test","task":"Visual Question Answering (VQA)","dataset":"VCR (Q-AR) test","model":"ERNIE-ViL-large(ensemble of 15 models)","rank_in_archive_order":2,"of":7,"metrics":{"Accuracy":"70.5"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-vcr-qa-r-test","task":"Visual Question Answering (VQA)","dataset":"VCR (QA-R) test","model":"ERNIE-ViL-large(ensemble of 15 models)","rank_in_archive_order":2,"of":8,"metrics":{"Accuracy":"86.1"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-std","task":"Visual Question Answering (VQA)","dataset":"VQA v2 test-std","model":"ERNIE-ViL-single model","rank_in_archive_order":16,"of":38,"metrics":{"number":"56.79","other":"65.24","overall":"74.93","yes/no":"90.83"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2006.16934","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}