Papers › VinVL+L: Enriching Visual Representation with Location Context in VQA
VinVL+L: Enriching Visual Representation with Location Context in VQA
Jiří Vyskočil, Lukáš Picek
In this paper, we describe a novel method - VinVL+L - that enriches the visual representations (i.e. object tags and region features) of the State-of-the-Art Vision and Language (VL) method - VinVL - with Location information. To verify the importance of such metadata for VL models, we (i) trained a Swin-B model on the Places365 dataset and obtained additional sets of visual and tag features; both were made public to allow reproducibility and further experiments, (ii) did an architectural update to the existing VinVL method to include the new feature sets, and (iii) provide a qualitative and quantitative evaluation. By including just binary location metadata, the VinVL+L method provides incremental improvement to the State-of-the-Art VinVL in Visual Question Answering (VQA). The VinVL+L achieved an accuracy of 64.85% and increased the performance by +0.32% in terms of accuracy on the GQA dataset; the statistical significance of the new representations is verified via Approximate Randomization. The code and newly generated sets of features are available at https://github.com/vyskocj/VinVL-L.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Question Answering (VQA) | GQA Test2019 | VinVL+L | Accuracy | 64.85 | #10 of 127 | Archive leaderboard | report |
| Visual Question Answering (VQA) | GQA Test2019 | VinVL+L | Binary | 82.59 | #10 of 127 | Archive leaderboard | report |
| Visual Question Answering (VQA) | GQA Test2019 | VinVL+L | Consistency | 94.0 | #10 of 127 | Archive leaderboard | report |
| Visual Question Answering (VQA) | GQA Test2019 | VinVL+L | Distribution | 4.59 | #10 of 127 | Archive leaderboard | report |
| Visual Question Answering (VQA) | GQA Test2019 | VinVL+L | Open | 49.19 | #10 of 127 | Archive leaderboard | report |
| Visual Question Answering (VQA) | GQA Test2019 | VinVL+L | Plausibility | 84.91 | #10 of 127 | Archive leaderboard | report |
| Visual Question Answering (VQA) | GQA Test2019 | VinVL+L | Validity | 96.62 | #10 of 127 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections