Papers › VinVL+L: Enriching Visual Representation with Location Context in VQA

VinVL+L: Enriching Visual Representation with Location Context in VQA

22 Feb 2023Computer Vision Winter Workshop 2023 2archive 2025-07-28

Jiří Vyskočil, Lukáš Picek

In this paper, we describe a novel method - VinVL+L - that enriches the visual representations (i.e. object tags and region features) of the State-of-the-Art Vision and Language (VL) method - VinVL - with Location information. To verify the importance of such metadata for VL models, we (i) trained a Swin-B model on the Places365 dataset and obtained additional sets of visual and tag features; both were made public to allow reproducibility and further experiments, (ii) did an architectural update to the existing VinVL method to include the new feature sets, and (iii) provide a qualitative and quantitative evaluation. By including just binary location metadata, the VinVL+L method provides incremental improvement to the State-of-the-Art VinVL in Visual Question Answering (VQA). The VinVL+L achieved an accuracy of 64.85% and increased the performance by +0.32% in terms of accuracy on the GQA dataset; the statistical significance of the new representations is verified via Approximate Randomization. The code and newly generated sets of features are available at https://github.com/vyskocj/VinVL-L.

PaperPDFCode

Code

vyskocj/VinVL-L mentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringTAGVisual Question AnsweringVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering (VQA) GQA Test2019 VinVL+L Accuracy 64.85 #10 of 127 Archive leaderboard report
Visual Question Answering (VQA) GQA Test2019 VinVL+L Binary 82.59 #10 of 127 Archive leaderboard report
Visual Question Answering (VQA) GQA Test2019 VinVL+L Consistency 94.0 #10 of 127 Archive leaderboard report
Visual Question Answering (VQA) GQA Test2019 VinVL+L Distribution 4.59 #10 of 127 Archive leaderboard report
Visual Question Answering (VQA) GQA Test2019 VinVL+L Open 49.19 #10 of 127 Archive leaderboard report
Visual Question Answering (VQA) GQA Test2019 VinVL+L Plausibility 84.91 #10 of 127 Archive leaderboard report
Visual Question Answering (VQA) GQA Test2019 VinVL+L Validity 96.62 #10 of 127 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections