{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/redit-re-evaluating-large-visual-question","title":"ReDiT: Re‑evaluating large visual question answering model confidence by defining input scenario Difficulty and applying Temperature mapping","arxiv_id":null,"date":"2025-01-06","proceeding":"Multimedia Systems 2025 1","authors":["Modafar Al-Shouha","Gábor Szűcs"],"abstract":"Large models (LMs) have achieved remarkable results in vision-language tasks. Such models are trained on vast amount of data then fine-tuned for downstream tasks like visual question answering (VQA). This wide exposure to data along with the complexity in a multi-modal setup (e.g. VQA) demand formalizing an extended definition of what constitutes out-of-distribution (OOD) condition for these models. Moreover, the input difficulty is expected to influence the model's performance, and it should be reflected on its confidence scoring. In this work, we primarily address large visual question answering (LVQA) models. We extend the classical boundaries of OOD definition and introduce a novel customizable dataset that simulates various challenges for LVQA models; i.e. 3U-VQA dataset. Moreover, we present a categorical scale to assess the input scenario difficulty. This scale is used to improve the reliability of the answer confidence score by re-evaluating it through adjusting a temperature parameter in the softmax function. Lastly, we study the credibility of our categorization and show that our re-evaluating method assists in reducing the overlap between correct and incorrect LVQA model predictions' scores.","url_abs":"https://doi.org/10.1007/s00530-024-01629-w","url_pdf":"https://trebuchet.public.springernature.app/get_content/d37db04d-0d30-4fba-9e61-edc7e190dc0d","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"redit-re-evaluating-large-visual-question","repo_url":"https://github.com/modafarshouha/ReDiT","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[{"slug":"3u-vqa","name":"3U-VQA","full_name":"Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}