Papers › ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

30 Jun 2020arXiv:2006.16934archive 2025-07-28

Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic connections (objects, attributes of objects and relationships between objects) across vision and language, which are essential to vision-language cross-modal tasks. Utilizing scene graphs of visual scenes, ERNIE-ViL constructs Scene Graph Prediction tasks, i.e., Object Prediction, Attribute Prediction and Relationship Prediction tasks in the pre-training phase. Specifically, these prediction tasks are implemented by predicting nodes of different types in the scene graph parsed from the sentence. Thus, ERNIE-ViL can learn the joint representations characterizing the alignments of the detailed semantics across vision and language. After pre-training on large scale image-text aligned datasets, we validate the effectiveness of ERNIE-ViL on 5 cross-modal downstream tasks. ERNIE-ViL achieves state-of-the-art performances on all these tasks and ranks the first place on the VCR leaderboard with an absolute improvement of 3.7%.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributePredictionReferring Expression ComprehensionSentenceVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering (VQA) VCR (Q-A) test ERNIE-ViL-large(ensemble of 15 models) Accuracy 81.6 #2 of 11 Archive leaderboard report
Visual Question Answering (VQA) VCR (Q-AR) test ERNIE-ViL-large(ensemble of 15 models) Accuracy 70.5 #2 of 7 Archive leaderboard report
Visual Question Answering (VQA) VCR (QA-R) test ERNIE-ViL-large(ensemble of 15 models) Accuracy 86.1 #2 of 8 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std ERNIE-ViL-single model number 56.79 #16 of 38 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std ERNIE-ViL-single model other 65.24 #16 of 38 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std ERNIE-ViL-single model overall 74.93 #16 of 38 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std ERNIE-ViL-single model yes/no 90.83 #16 of 38 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections