Papers › QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning

QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning

6 May 2022Findings (NAACL) 2022 7arXiv:2205.03075archive 2025-07-28

Zechen Li, Anders Søgaard

Synthetic datasets have successfully been used to probe visual question-answering datasets for their reasoning abilities. CLEVR (johnson2017clevr), for example, tests a range of visual reasoning abilities. The questions in CLEVR focus on comparisons of shapes, colors, and sizes, numerical reasoning, and existence claims. This paper introduces a minimally biased, diagnostic visual question-answering dataset, QLEVR, that goes beyond existential and numerical quantification and focus on more complex quantifiers and their combinations, e.g., asking whether there are more than two red balls that are smaller than at least three blue balls in an image. We describe how the dataset was created and present a first evaluation of state-of-the-art visual question-answering models, showing that QLEVR presents a formidable challenge to our current models. Code and Dataset are available at https://github.com/zechenli03/QLEVR

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zechenli03/qlevr officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DiagnosticQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Datasets

Introduced by this paper, per the archive.

QLEVR

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering (VQA) QLEVR MAC Overall Accuracy 66.5 #1 of 5 Archive leaderboard report
Visual Question Answering (VQA) QLEVR CNN+LSTM Overall Accuracy 65.9 #2 of 5 Archive leaderboard report
Visual Question Answering (VQA) QLEVR BERT Overall Accuracy 65.8 #3 of 5 Archive leaderboard report
Visual Question Answering (VQA) QLEVR LSTM Overall Accuracy 64.6 #4 of 5 Archive leaderboard report
Visual Question Answering (VQA) QLEVR Q-type Overall Accuracy 50.0 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections