Papers › ChiQA: A Large Scale Image-based Real-World Question Answering Dataset for Multi-Modal...

ChiQA: A Large Scale Image-based Real-World Question Answering Dataset for Multi-Modal Understanding

5 Aug 2022arXiv:2208.03030archive 2025-07-28

Bingning Wang, Feiyang Lv, Ting Yao, Yiming Yuan, Jin Ma, Yu Luo, Haijin Liang

Visual question answering is an important task in both natural language and vision understanding. However, in most of the public visual question answering datasets such as VQA, CLEVR, the questions are human generated that specific to the given image, such as `What color are her eyes?'. The human generated crowdsourcing questions are relatively simple and sometimes have the bias toward certain entities or attributes. In this paper, we introduce a new question answering dataset based on image-ChiQA. It contains the real-world queries issued by internet users, combined with several related open-domain images. The system should determine whether the image could answer the question or not. Different from previous VQA datasets, the questions are real-world image-independent queries that are more various and unbiased. Compared with previous image-retrieval or image-caption datasets, the ChiQA not only measures the relatedness but also measures the answerability, which demands more fine-grained vision and language reasoning. ChiQA contains more than 40K questions and more than 200K question-images pairs. A three-level 2/1/0 label is assigned to each pair indicating perfect answer, partially answer and irrelevant. Data analysis shows ChiQA requires a deep understanding of both language and vision, including grounding, comparisons, and reading. We evaluate several state-of-the-art visual-language models such as ALBEF, demonstrating that there is still a large room for improvements on ChiQA.

PaperPDFCode

Code

benywon/ChiQA officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image RetrievalQuestion AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Datasets

Introduced by this paper, per the archive.

ChiQA

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

ALBEF

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections