{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/can-i-trust-your-answer-visually-grounded","title":"Can I Trust Your Answer? Visually Grounded Video Question Answering","arxiv_id":"2309.01327","date":"2023-09-04","proceeding":"CVPR 2024 1","authors":["Junbin Xiao","Angela Yao","Yicong Li","Tat Seng Chua"],"abstract":"We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video-language understanding. Specifically, by forcing vision-language models (VLMs) to answer questions and simultaneously provide visual evidence, we seek to ascertain the extent to which the predictions of such techniques are genuinely anchored in relevant video content, versus spurious correlations from language or irrelevant visual context. Towards this, we construct NExT-GQA -- an extension of NExT-QA with 10.5$K$ temporal grounding (or location) labels tied to the original QA pairs. With NExT-GQA, we scrutinize a series of state-of-the-art VLMs. Through post-hoc attention analysis, we find that these models are extremely weak in substantiating the answers despite their strong QA performance. This exposes the limitation of current VLMs in making reliable predictions. As a remedy, we further explore and propose a grounded-QA method via Gaussian mask optimization and cross-modal learning. Experiments with different backbones demonstrate that this grounding mechanism improves both grounding and QA. With these efforts, we aim to push towards trustworthy VLMs in VQA systems. Our dataset and code are available at https://github.com/doc-doc/NExT-GQA.","url_abs":"https://arxiv.org/abs/2309.01327v2","url_pdf":"https://arxiv.org/pdf/2309.01327v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"can-i-trust-your-answer-visually-grounded","repo_url":"https://github.com/doc-doc/next-gqa","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"grounded-video-question-answering","task_name":"Grounded Video Question Answering"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-grounding","task_name":"Video Grounding"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[{"slug":"next-gqa","name":"NExT-GQA","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2309.01327","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2309.01327"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/doc-doc/next-gqa","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_honours":1,"ran":2,"ran_violates":1,"ran_fixture":1,"unverified":2},"by_repo_kind":{"official":{"samples":7,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"0464dc8467b456dc","entry":"build_relative_position","repo":"doc-doc/next-gqa","repo_kind":"official","path":"code/FrozenGQA/model/deberta.py","file_url":"https://github.com/doc-doc/next-gqa/blob/HEAD/code/FrozenGQA/model/deberta.py","link_basis":"plan_row","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0464dc8467b456dc"}},{"code_sha256_prefix":"03e99c761c545617","entry":"duplicate_interleave","repo":"doc-doc/next-gqa","repo_kind":"official","path":"code/FrozenGQA/model/gptj.py","file_url":"https://github.com/doc-doc/next-gqa/blob/HEAD/code/FrozenGQA/model/gptj.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"03e99c761c545617"}},{"code_sha256_prefix":"7ea21b9da7ef5be6","entry":"fixed_pos_embedding","repo":"doc-doc/next-gqa","repo_kind":"official","path":"code/FrozenGQA/model/gptj.py","file_url":"https://github.com/doc-doc/next-gqa/blob/HEAD/code/FrozenGQA/model/gptj.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7ea21b9da7ef5be6"}},{"code_sha256_prefix":"567c8b010f8e20d2","entry":"make_log_bucket_position","repo":"doc-doc/next-gqa","repo_kind":"official","path":"code/FrozenGQA/model/deberta.py","file_url":"https://github.com/doc-doc/next-gqa/blob/HEAD/code/FrozenGQA/model/deberta.py","link_basis":"plan_row","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"567c8b010f8e20d2"}},{"code_sha256_prefix":"c66149010337c505","entry":"rotate_every_two","repo":"doc-doc/next-gqa","repo_kind":"official","path":"code/FrozenGQA/model/gptj.py","file_url":"https://github.com/doc-doc/next-gqa/blob/HEAD/code/FrozenGQA/model/gptj.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c66149010337c505"}},{"code_sha256_prefix":"05732f1a5126672d","entry":"get_mask","repo":"doc-doc/next-gqa","repo_kind":"official","path":"code/FrozenGQA/model/deberta.py","file_url":"https://github.com/doc-doc/next-gqa/blob/HEAD/code/FrozenGQA/model/deberta.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"05732f1a5126672d"}},{"code_sha256_prefix":"d8f309037de819b0","entry":"load_tf_weights_in_gpt_neo","repo":"doc-doc/next-gqa","repo_kind":"official","path":"code/FrozenGQA/model/gptneo.py","file_url":"https://github.com/doc-doc/next-gqa/blob/HEAD/code/FrozenGQA/model/gptneo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d8f309037de819b0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}