Papers › LXMERT Model Compression for Visual Question Answering

LXMERT Model Compression for Visual Question Answering

23 Oct 2023arXiv:2310.15325archive 2025-07-28

Maryam Hashemi, Ghazaleh Mahmoudi, Sara Kodeiri, Hadi Sheikhi, Sauleh Eetemadi

Large-scale pretrained models such as LXMERT are becoming popular for learning cross-modal representations on text-image pairs for vision-language tasks. According to the lottery ticket hypothesis, NLP and computer vision models contain smaller subnetworks capable of being trained in isolation to full performance. In this paper, we combine these observations to evaluate whether such trainable subnetworks exist in LXMERT when fine-tuned on the VQA task. In addition, we perform a model size cost-benefit analysis by investigating how much pruning can be done without significant loss in accuracy. Our experiment results demonstrate that LXMERT can be effectively pruned by 40%-60% in size with 3% loss in accuracy.

PaperPDFCode

Code

ghazaleh-mahmoodi/lxmert_compression officialmentioned on GitHubpytorch report
pwc-1/Paper-9 mindspore report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Model CompressionVisual Question AnsweringVisual Question Answering (VQA)model

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering VQA v2 test-dev LXMERT (low-magnitude pruning) Accuracy 70.72 #10 of 11 Archive leaderboard report
Visual Question Answering VQA v2 test-std LXMERT (low-magnitude pruning) Accuracy 70.87 #1 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

LXMERTPruning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections