Papers › Towards VQA Models That Can Read

Towards VQA Models That Can Read

18 Apr 2019CVPR 2019 6arXiv:1904.08920archive 2025-07-28

Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, Marcus Rohrbach

Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to facilitate progress on this important problem. Existing datasets either have a small proportion of questions about text (e.g., the VQA dataset) or are too small (e.g., the VizWiz dataset). TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Second, we introduce a novel model architecture that reads text in the image, reasons about it in the context of the image and the question, and predicts an answer which might be a deduction based on the text and the image or composed of the strings found in the image. Consequently, we call our approach Look, Read, Reason & Answer (LoRRA). We show that LoRRA outperforms existing state-of-the-art VQA models on our TextVQA dataset. We find that the gap between human performance and machine performance is significantly larger on TextVQA than on VQA 2.0, suggesting that TextVQA is well-suited to benchmark progress along directions complementary to VQA 2.0.

PaperPDFConference PDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

facebookresearch/pythia officialmentioned in papermentioned on GitHubpytorchNOASSERTION report
ZephyrZhuQi/ssbaseline mentioned on GitHubpytorchNOASSERTION report
allenai/pythia mentioned on GitHubpytorchNOASSERTION report
facebookresearch/mmf mentioned on GitHubpytorchNOASSERTION report
jackroos/pythia mentioned on GitHubpytorchNOASSERTION report
ronghanghu/pythia mentioned on GitHubpytorchNOASSERTION report
zwxalgorithm/pythia mentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Visual Question Answering (VQA)

1 archive task tag without a task page not shown.

Datasets

Introduced by this paper, per the archive.

TextVQA

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering (VQA) VQA v2 test-dev Pythia v0.3 + LoRRA Accuracy 69.21 #34 of 56 Archive leaderboard report
Visual Question Answering (VQA) VizWiz 2018 Pythia v0.3 overall 54.72 #3 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections