{"url":"/dataset/visual-question-answering-v2-0","name":"Visual Question Answering v2.0","full_name":"VQA v2.0","description_markdown":"Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images. These questions require an understanding of vision, language and commonsense knowledge to answer. It is the second version of the VQA dataset.\r\n\r\n- 265,016 images (COCO and abstract scenes)\r\n- At least 3 questions (5.4 questions on average) per image\r\n- 10 ground truth answers per question\r\n- 3 plausible (but likely incorrect) answers per question\r\n- Automatic evaluation metric\r\n\r\nThe [first version of the dataset](/dataset/visual-question-answering) was released in October 2015.","description_withheld":null,"homepage":"https://visualqa.org/","introduced_date":"2016-12-02","introduced_date_note":null,"introduced_by":{"paper":"/paper/making-the-v-in-vqa-matter-elevating-the-role","title":"Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering","first_author":"Yash Goyal","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Visual Question Answering","url":"/task/visual-question-answering-1","datasets_with_task":"/datasets/task/visual-question-answering-1"}],"languages":[],"variants":["VQA 2.0","VQA v2","VQA v2 test-dev","VQA v2 test-std","Visual Question Answering v2.0","VQA v2 val"],"data_loaders":[{"repo":"https://github.com/facebookresearch/ParlAI","url":"https://parl.ai/docs/tasks.html#vqav2","frameworks":["pytorch"]},{"repo":"https://github.com/allenai/allennlp-models","url":"https://docs.allennlp.org/models/main/models/vision/dataset_readers/vqav2/","frameworks":["pytorch"]}],"num_papers_in_archive":366,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-dev","task":"Visual Question Answering (VQA)","dataset_variant":"VQA v2 test-dev","rows":56,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"PaLI","paper":"/paper/pali-a-jointly-scaled-multilingual-language","metrics":{"Accuracy":"84.3"},"code_links":[{"title":"google-research/big_vision","url":"https://github.com/google-research/big_vision"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-std","task":"Visual Question Answering (VQA)","dataset_variant":"VQA v2 test-std","rows":38,"metrics":["overall","yes/no","number","other"],"first_row_in_archive_order":{"model":"BEiT-3","paper":"/paper/image-as-a-foreign-language-beit-pretraining","metrics":{"overall":"84.03"},"code_links":[{"title":"microsoft/unilm","url":"https://github.com/microsoft/unilm/tree/master/beit"},{"title":"lyan62/data-curation","url":"https://github.com/lyan62/data-curation"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-dev-1","task":"Visual Question Answering","dataset_variant":"VQA v2 test-dev","rows":11,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"BLIP-2 ViT-G OPT 6.7B (fine-tuned)","paper":"/paper/blip-2-bootstrapping-language-image-pre","metrics":{"Accuracy":"82.30"},"code_links":[{"title":"huggingface/transformers","url":"https://github.com/huggingface/transformers"},{"title":"salesforce/lavis","url":"https://github.com/salesforce/lavis"},{"title":"thudm/visualglm-6b","url":"https://github.com/thudm/visualglm-6b"},{"title":"baaivision/eva","url":"https://github.com/baaivision/eva"},{"title":"junshutang/Make-It-3D","url":"https://github.com/junshutang/Make-It-3D"},{"title":"facebookresearch/multimodal","url":"https://github.com/facebookresearch/multimodal"},{"title":"unispac/visual-adversarial-examples-jailbreak-large-language-models","url":"https://github.com/unispac/visual-adversarial-examples-jailbreak-large-language-models"},{"title":"yukw777/videoblip","url":"https://github.com/yukw777/videoblip"},{"title":"alibaba/graphtranslator","url":"https://github.com/alibaba/graphtranslator"},{"title":"gregor-ge/mblip","url":"https://github.com/gregor-ge/mblip"},{"title":"linzhiqiu/clip-flant5","url":"https://github.com/linzhiqiu/clip-flant5"},{"title":"kdr/videorag-mrr2024","url":"https://github.com/kdr/videorag-mrr2024"},{"title":"rabiulcste/vqazero","url":"https://github.com/rabiulcste/vqazero"},{"title":"jiwanchung/vlis","url":"https://github.com/jiwanchung/vlis"},{"title":"yangyucheng000/University","url":"https://github.com/yangyucheng000/University/tree/main/model-2/blip_2"},{"title":"2024-MindSpore-1/Code2","url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/blip_2"},{"title":"albertotestoni/ndq_visual_objects","url":"https://github.com/albertotestoni/ndq_visual_objects"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-val","task":"Visual Question Answering (VQA)","dataset_variant":"VQA v2 val","rows":11,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"BLIP-2 ViT-G FlanT5 XXL (zero-shot)","paper":"/paper/blip-2-bootstrapping-language-image-pre","metrics":{"Accuracy":"65.2"},"code_links":[{"title":"huggingface/transformers","url":"https://github.com/huggingface/transformers"},{"title":"salesforce/lavis","url":"https://github.com/salesforce/lavis"},{"title":"thudm/visualglm-6b","url":"https://github.com/thudm/visualglm-6b"},{"title":"baaivision/eva","url":"https://github.com/baaivision/eva"},{"title":"junshutang/Make-It-3D","url":"https://github.com/junshutang/Make-It-3D"},{"title":"facebookresearch/multimodal","url":"https://github.com/facebookresearch/multimodal"},{"title":"unispac/visual-adversarial-examples-jailbreak-large-language-models","url":"https://github.com/unispac/visual-adversarial-examples-jailbreak-large-language-models"},{"title":"yukw777/videoblip","url":"https://github.com/yukw777/videoblip"},{"title":"alibaba/graphtranslator","url":"https://github.com/alibaba/graphtranslator"},{"title":"gregor-ge/mblip","url":"https://github.com/gregor-ge/mblip"},{"title":"linzhiqiu/clip-flant5","url":"https://github.com/linzhiqiu/clip-flant5"},{"title":"kdr/videorag-mrr2024","url":"https://github.com/kdr/videorag-mrr2024"},{"title":"rabiulcste/vqazero","url":"https://github.com/rabiulcste/vqazero"},{"title":"jiwanchung/vlis","url":"https://github.com/jiwanchung/vlis"},{"title":"yangyucheng000/University","url":"https://github.com/yangyucheng000/University/tree/main/model-2/blip_2"},{"title":"2024-MindSpore-1/Code2","url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/blip_2"},{"title":"albertotestoni/ndq_visual_objects","url":"https://github.com/albertotestoni/ndq_visual_objects"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-val-1","task":"Visual Question Answering","dataset_variant":"VQA v2 val","rows":4,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"BLIP-2 ViT-G OPT 6.7B (fine-tuned)","paper":"/paper/blip-2-bootstrapping-language-image-pre","metrics":{"Accuracy":"82.19"},"code_links":[{"title":"huggingface/transformers","url":"https://github.com/huggingface/transformers"},{"title":"salesforce/lavis","url":"https://github.com/salesforce/lavis"},{"title":"thudm/visualglm-6b","url":"https://github.com/thudm/visualglm-6b"},{"title":"baaivision/eva","url":"https://github.com/baaivision/eva"},{"title":"junshutang/Make-It-3D","url":"https://github.com/junshutang/Make-It-3D"},{"title":"facebookresearch/multimodal","url":"https://github.com/facebookresearch/multimodal"},{"title":"unispac/visual-adversarial-examples-jailbreak-large-language-models","url":"https://github.com/unispac/visual-adversarial-examples-jailbreak-large-language-models"},{"title":"yukw777/videoblip","url":"https://github.com/yukw777/videoblip"},{"title":"alibaba/graphtranslator","url":"https://github.com/alibaba/graphtranslator"},{"title":"gregor-ge/mblip","url":"https://github.com/gregor-ge/mblip"},{"title":"linzhiqiu/clip-flant5","url":"https://github.com/linzhiqiu/clip-flant5"},{"title":"kdr/videorag-mrr2024","url":"https://github.com/kdr/videorag-mrr2024"},{"title":"rabiulcste/vqazero","url":"https://github.com/rabiulcste/vqazero"},{"title":"jiwanchung/vlis","url":"https://github.com/jiwanchung/vlis"},{"title":"yangyucheng000/University","url":"https://github.com/yangyucheng000/University/tree/main/model-2/blip_2"},{"title":"2024-MindSpore-1/Code2","url":"https://github.com/2024-MindSpore-1/Code2/tree/main/model-1/blip_2"},{"title":"albertotestoni/ndq_visual_objects","url":"https://github.com/albertotestoni/ndq_visual_objects"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-std-1","task":"Visual Question Answering","dataset_variant":"VQA v2 test-std","rows":3,"metrics":["Accuracy","number","other","overall","yes/no"],"first_row_in_archive_order":{"model":"LXMERT (low-magnitude pruning)","paper":"/paper/lxmert-model-compression-for-visual-question","metrics":{"Accuracy":"70.87"},"code_links":[{"title":"ghazaleh-mahmoodi/lxmert_compression","url":"https://github.com/ghazaleh-mahmoodi/lxmert_compression"},{"title":"pwc-1/Paper-9","url":"https://github.com/pwc-1/Paper-9/tree/main/lxmert"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-1","task":"Visual Question Answering","dataset_variant":"VQA v2","rows":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"RLHF-V","paper":"/paper/rlhf-v-towards-trustworthy-mllms-via-behavior","metrics":{"Accuracy":"80"},"code_links":[{"title":"openbmb/minicpm-v","url":"https://github.com/openbmb/minicpm-v"},{"title":"rlhf-v/rlhf-v","url":"https://github.com/rlhf-v/rlhf-v"},{"title":"tidedra/vl-rlhf","url":"https://github.com/tidedra/vl-rlhf"},{"title":"exgc/r1v-free","url":"https://github.com/exgc/r1v-free"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/cumo-scaling-multimodal-llm-with-co-upcycled","title":"CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts","date":"2024-05-09","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":10,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-to-localize-objects-improves-spatial","title":"Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs","date":"2024-04-11","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/internvl-scaling-up-vision-foundation-models","title":"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks","date":"2023-12-21","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/lyrics-boosting-fine-grained-language-vision","title":"Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects","date":"2023-12-08","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/rlhf-v-towards-trustworthy-mllms-via-behavior","title":"RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback","date":"2023-12-01","rows_on_this_dataset":1,"code_links":4,"syntology":null},{"paper":"/paper/lxmert-model-compression-for-visual-question","title":"LXMERT Model Compression for Visual Question Answering","date":"2023-10-23","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/implicit-differentiable-outlier-detection","title":"Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis","date":"2023-09-21","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/generative-pretraining-in-multimodality","title":"Emu: Generative Pretraining in Multimodality","date":"2023-07-11","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":0,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/one-peace-exploring-one-general","title":"ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities","date":"2023-05-18","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":2,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/valor-vision-audio-language-omni-perception","title":"VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset","date":"2023-04-17","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/prismer-a-vision-language-model-with-an","title":"Prismer: A Vision-Language Model with Multi-Task Experts","date":"2023-03-04","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/differentiable-outlier-detection-enable","title":"Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis","date":"2023-02-11","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mplug-2-a-modularized-multi-modal-foundation","title":"mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video","date":"2023-02-01","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":19,"samples_ran":9,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/blip-2-bootstrapping-language-image-pre","title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","date":"2023-01-30","rows_on_this_dataset":18,"code_links":17,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":4,"samples_unverified":4,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/toward-building-general-foundation-models-for","title":"Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks","date":"2023-01-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/x-2-vlm-all-in-one-pre-trained-model-for","title":"X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks","date":"2022-11-22","rows_on_this_dataset":4,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":2,"samples_unverified":4,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/plug-and-play-vqa-zero-shot-vqa-by-conjoining","title":"Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training","date":"2022-10-17","rows_on_this_dataset":2,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/pali-a-jointly-scaled-multilingual-language","title":"PaLI: A Jointly-Scaled Multilingual Language-Image Model","date":"2022-09-14","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/image-as-a-foreign-language-beit-pretraining","title":"Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks","date":"2022-08-22","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/prompt-tuning-for-generative-multimodal","title":"Prompt Tuning for Generative Multimodal Pretrained Models","date":"2022-08-04","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/lako-knowledge-driven-visual-question","title":"LaKo: Knowledge-driven Visual Question Answering via Late Knowledge-to-Text Injection","date":"2022-07-26","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/language-models-are-general-purpose","title":"Language Models are General-Purpose Interfaces","date":"2022-06-13","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mplug-effective-and-efficient-vision-language","title":"mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections","date":"2022-05-24","rows_on_this_dataset":2,"code_links":3,"syntology":null},{"paper":"/paper/coca-contrastive-captioners-are-image-text","title":"CoCa: Contrastive Captioners are Image-Text Foundation Models","date":"2022-05-04","rows_on_this_dataset":1,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":9,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/flamingo-a-visual-language-model-for-few-shot-1","title":"Flamingo: a Visual Language Model for Few-Shot Learning","date":"2022-04-29","rows_on_this_dataset":3,"code_links":5,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":24,"samples_ran":18,"samples_unverified":6,"pointer_only_for_licence":7,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unifying-architectures-tasks-and-modalities","title":"OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework","date":"2022-02-07","rows_on_this_dataset":2,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/florence-a-new-foundation-model-for-computer","title":"Florence: A New Foundation Model for Computer Vision","date":"2021-11-22","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/achieving-human-parity-on-visual-question","title":"Achieving Human Parity on Visual Question Answering","date":"2021-11-17","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/multi-grained-vision-language-pre-training","title":"Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts","date":"2021-11-16","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/enabling-multimodal-generation-on-clip-via","title":"Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation","date":"2021-11-16","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/vlmo-unified-vision-language-pre-training","title":"VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts","date":"2021-11-03","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/a-good-prompt-is-worth-millions-of-parameters","title":"A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models","date":"2021-10-16","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/coarse-to-fine-reasoning-for-visual-question","title":"Coarse-to-Fine Reasoning for Visual Question Answering","date":"2021-10-06","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":4,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/simvlm-simple-visual-language-model","title":"SimVLM: Simple Visual Language Model Pretraining with Weak Supervision","date":"2021-08-24","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":37,"samples_ran":18,"samples_unverified":19,"pointer_only_for_licence":28,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/align-before-fuse-vision-and-language","title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","date":"2021-07-16","rows_on_this_dataset":2,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":3,"samples_unverified":2,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/multimodal-few-shot-learning-with-frozen","title":"Multimodal Few-Shot Learning with Frozen Language Models","date":"2021-06-25","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/vilt-vision-and-language-transformer-without","title":"ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision","date":"2021-02-05","rows_on_this_dataset":1,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":1,"samples_unverified":3,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/vinvl-making-visual-representations-matter-in","title":"VinVL: Revisiting Visual Representations in Vision-Language Models","date":"2021-01-02","rows_on_this_dataset":2,"code_links":7,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/ernie-vil-knowledge-enhanced-vision-language","title":"ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph","date":"2020-06-30","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/sparse-and-continuous-attention-mechanisms","title":"Sparse and Continuous Attention Mechanisms","date":"2020-06-12","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":15,"samples_ran":12,"samples_unverified":3,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/deep-multimodal-neural-architecture-search","title":"Deep Multimodal Neural Architecture Search","date":"2020-04-25","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/oscar-object-semantics-aligned-pre-training","title":"Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks","date":"2020-04-13","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":23,"samples_ran":11,"samples_unverified":12,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/visual-commonsense-r-cnn","title":"Visual Commonsense R-CNN","date":"2020-02-27","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":1,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/in-defense-of-grid-features-for-visual","title":"In Defense of Grid Features for Visual Question Answering","date":"2020-01-10","rows_on_this_dataset":3,"code_links":2,"syntology":null},{"paper":"/paper/compact-trilinear-interaction-for-visual","title":"Compact Trilinear Interaction for Visual Question Answering","date":"2019-09-26","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":1,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/uniter-learning-universal-image-text-1","title":"UNITER: UNiversal Image-TExt Representation Learning","date":"2019-09-25","rows_on_this_dataset":2,"code_links":7,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unified-vision-language-pre-training-for","title":"Unified Vision-Language Pre-Training for Image Captioning and VQA","date":"2019-09-24","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":14,"samples_ran":5,"samples_unverified":9,"pointer_only_for_licence":13,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/vl-bert-pre-training-of-generic-visual","title":"VL-BERT: Pre-training of Generic Visual-Linguistic Representations","date":"2019-08-22","rows_on_this_dataset":3,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/lxmert-learning-cross-modality-encoder","title":"LXMERT: Learning Cross-Modality Encoder Representations from Transformers","date":"2019-08-20","rows_on_this_dataset":2,"code_links":9,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":15,"samples_ran":4,"samples_unverified":11,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/visualbert-a-simple-and-performant-baseline","title":"VisualBERT: A Simple and Performant Baseline for Vision and Language","date":"2019-08-09","rows_on_this_dataset":2,"code_links":10,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":4,"samples_unverified":5,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/vilbert-pretraining-task-agnostic","title":"ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks","date":"2019-08-06","rows_on_this_dataset":1,"code_links":11,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":34,"samples_ran":10,"samples_unverified":24,"pointer_only_for_licence":34,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/graph-reasoning-networks-for-visual-question","title":"Bilinear Graph Networks for Visual Question Answering","date":"2019-07-23","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/deep-modular-co-attention-networks-for-visual-1","title":"Deep Modular Co-Attention Networks for Visual Question Answering","date":"2019-06-25","rows_on_this_dataset":2,"code_links":7,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/rubi-reducing-unimodal-biases-in-visual","title":"RUBi: Reducing Unimodal Biases in Visual Question Answering","date":"2019-06-24","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/190600513","title":"Generating Question Relevant Captions to Aid Visual Question Answering","date":"2019-06-03","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/towards-vqa-models-that-can-read","title":"Towards VQA Models That Can Read","date":"2019-04-18","rows_on_this_dataset":1,"code_links":7,"syntology":null},{"paper":"/paper/murel-multimodal-relational-reasoning-for","title":"MUREL: Multimodal Relational Reasoning for Visual Question Answering","date":"2019-02-25","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/block-bilinear-superdiagonal-fusion-for","title":"BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection","date":"2019-01-31","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":0,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/bilinear-attention-networks","title":"Bilinear Attention Networks","date":"2018-05-21","rows_on_this_dataset":2,"code_links":8,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":4,"samples_unverified":9,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-to-count-objects-in-natural-images","title":"Learning to Count Objects in Natural Images for Visual Question Answering","date":"2018-02-15","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/tips-and-tricks-for-visual-question-answering","title":"Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge","date":"2017-08-09","rows_on_this_dataset":2,"code_links":10,"syntology":null},{"paper":"/paper/bottom-up-and-top-down-attention-for-image","title":"Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering","date":"2017-07-25","rows_on_this_dataset":1,"code_links":65,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":9,"samples_unverified":0,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mutan-multimodal-tucker-fusion-for-visual","title":"MUTAN: Multimodal Tucker Fusion for Visual Question Answering","date":"2017-05-18","rows_on_this_dataset":2,"code_links":6,"syntology":null},{"paper":"/paper/learning-to-reason-end-to-end-module-networks","title":"Learning to Reason: End-to-End Module Networks for Visual Question Answering","date":"2017-04-18","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/making-the-v-in-vqa-matter-elevating-the-role","title":"Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering","date":"2016-12-02","rows_on_this_dataset":3,"code_links":7,"syntology":null},{"paper":"/paper/multimodal-compact-bilinear-pooling-for","title":"Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding","date":"2016-06-06","rows_on_this_dataset":1,"code_links":10,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":34,"samples_harvested":313,"samples_ran":156,"samples_unverified":157,"pointer_only_for_licence":120,"papers_with_no_sample_that_ran":4,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}