Papers › ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang
We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic connections (objects, attributes of objects and relationships between objects) across vision and language, which are essential to vision-language cross-modal tasks. Utilizing scene graphs of visual scenes, ERNIE-ViL constructs Scene Graph Prediction tasks, i.e., Object Prediction, Attribute Prediction and Relationship Prediction tasks in the pre-training phase. Specifically, these prediction tasks are implemented by predicting nodes of different types in the scene graph parsed from the sentence. Thus, ERNIE-ViL can learn the joint representations characterizing the alignments of the detailed semantics across vision and language. After pre-training on large scale image-text aligned datasets, we validate the effectiveness of ERNIE-ViL on 5 cross-modal downstream tasks. ERNIE-ViL achieves state-of-the-art performances on all these tasks and ranks the first place on the VCR leaderboard with an absolute improvement of 3.7%.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Question Answering (VQA) | VCR (Q-A) test | ERNIE-ViL-large(ensemble of 15 models) | Accuracy | 81.6 | #2 of 11 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VCR (Q-AR) test | ERNIE-ViL-large(ensemble of 15 models) | Accuracy | 70.5 | #2 of 7 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VCR (QA-R) test | ERNIE-ViL-large(ensemble of 15 models) | Accuracy | 86.1 | #2 of 8 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VQA v2 test-std | ERNIE-ViL-single model | number | 56.79 | #16 of 38 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VQA v2 test-std | ERNIE-ViL-single model | other | 65.24 | #16 of 38 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VQA v2 test-std | ERNIE-ViL-single model | overall | 74.93 | #16 of 38 | Archive leaderboard | report |
| Visual Question Answering (VQA) | VQA v2 test-std | ERNIE-ViL-single model | yes/no | 90.83 | #16 of 38 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections