Papers › Simple Token-Level Confidence Improves Caption Correctness

Simple Token-Level Confidence Improves Caption Correctness

11 May 2023arXiv:2305.07021archive 2025-07-28

Suzanne Petryk, Spencer Whitehead, Joseph E. Gonzalez, Trevor Darrell, Anna Rohrbach, Marcus Rohrbach

The ability to judge whether a caption correctly describes an image is a critical part of vision-language understanding. However, state-of-the-art models often misinterpret the correctness of fine-grained details, leading to errors in outputs such as hallucinating objects in generated captions or poor compositional reasoning. In this work, we explore Token-Level Confidence, or TLC, as a simple yet surprisingly effective method to assess caption correctness. Specifically, we fine-tune a vision-language model on image captioning, input an image and proposed caption to the model, and aggregate either algebraic or learned token confidences over words or sequences to estimate image-caption consistency. Compared to sequence-level scores from pretrained models, TLC with algebraic confidence measures achieves a relative improvement in accuracy by 10% on verb understanding in SVO-Probes and outperforms prior state-of-the-art in image and group scores for compositional reasoning in Winoground by a relative 37% and 9%, respectively. When training data are available, a learned confidence estimator provides further improved performance, reducing object hallucination rates in MS COCO Captions by a relative 30% over the original model and setting a new state-of-the-art.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

HallucinationImage CaptioningLanguage ModellingObject HallucinationVisual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Reasoning Winoground OFA large (ITM) Group Score 7.25 #66 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (ITM) Image Score 10.25 #66 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (ITM) Text Score 30.75 #66 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (TLC-A) Group Score 17.50 #73 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (TLC-A) Image Score 27.00 #73 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (TLC-A) Text Score 29.25 #73 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA base (ITM) Group Score 6.50 #82 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA base (ITM) Image Score 10.75 #82 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA base (ITM) Text Score 26.75 #82 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA base (TLC-A) Group Score 13.75 #88 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA base (TLC-A) Image Score 23.50 #88 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA base (TLC-A) Text Score 24.50 #88 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA tiny (ITM) Group Score 4.50 #94 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA tiny (ITM) Image Score 7.75 #94 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA tiny (ITM) Text Score 22.75 #94 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA tiny (TLC-A) Group Score 6.75 #109 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA tiny (TLC-A) Image Score 15.75 #109 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA tiny (TLC-A) Text Score 16.50 #109 of 114 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

TLC

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections