{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/simple-token-level-confidence-improves","title":"Simple Token-Level Confidence Improves Caption Correctness","arxiv_id":"2305.07021","date":"2023-05-11","proceeding":null,"authors":["Suzanne Petryk","Spencer Whitehead","Joseph E. Gonzalez","Trevor Darrell","Anna Rohrbach","Marcus Rohrbach"],"abstract":"The ability to judge whether a caption correctly describes an image is a critical part of vision-language understanding. However, state-of-the-art models often misinterpret the correctness of fine-grained details, leading to errors in outputs such as hallucinating objects in generated captions or poor compositional reasoning. In this work, we explore Token-Level Confidence, or TLC, as a simple yet surprisingly effective method to assess caption correctness. Specifically, we fine-tune a vision-language model on image captioning, input an image and proposed caption to the model, and aggregate either algebraic or learned token confidences over words or sequences to estimate image-caption consistency. Compared to sequence-level scores from pretrained models, TLC with algebraic confidence measures achieves a relative improvement in accuracy by 10% on verb understanding in SVO-Probes and outperforms prior state-of-the-art in image and group scores for compositional reasoning in Winoground by a relative 37% and 9%, respectively. When training data are available, a learned confidence estimator provides further improved performance, reducing object hallucination rates in MS COCO Captions by a relative 30% over the original model and setting a new state-of-the-art.","url_abs":"https://arxiv.org/abs/2305.07021v1","url_pdf":"https://arxiv.org/pdf/2305.07021v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"object-hallucination","task_name":"Object Hallucination"},{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[{"method_slug":"tlc","method_name":"TLC"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"OFA large (ITM)","rank_in_archive_order":66,"of":114,"metrics":{"Group Score":"7.25","Image Score":"10.25","Text Score":"30.75"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"OFA large (TLC-A)","rank_in_archive_order":73,"of":114,"metrics":{"Group Score":"17.50","Image Score":"27.00","Text Score":"29.25"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"OFA base (ITM)","rank_in_archive_order":82,"of":114,"metrics":{"Group Score":"6.50","Image Score":"10.75","Text Score":"26.75"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"OFA base (TLC-A)","rank_in_archive_order":88,"of":114,"metrics":{"Group Score":"13.75","Image Score":"23.50","Text Score":"24.50"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"OFA tiny (ITM)","rank_in_archive_order":94,"of":114,"metrics":{"Group Score":"4.50","Image Score":"7.75","Text Score":"22.75"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"OFA tiny (TLC-A)","rank_in_archive_order":109,"of":114,"metrics":{"Group Score":"6.75","Image Score":"15.75","Text Score":"16.50"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2305.07021","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}