Papers › Simple Token-Level Confidence Improves Caption Correctness
Simple Token-Level Confidence Improves Caption Correctness
Suzanne Petryk, Spencer Whitehead, Joseph E. Gonzalez, Trevor Darrell, Anna Rohrbach, Marcus Rohrbach
The ability to judge whether a caption correctly describes an image is a critical part of vision-language understanding. However, state-of-the-art models often misinterpret the correctness of fine-grained details, leading to errors in outputs such as hallucinating objects in generated captions or poor compositional reasoning. In this work, we explore Token-Level Confidence, or TLC, as a simple yet surprisingly effective method to assess caption correctness. Specifically, we fine-tune a vision-language model on image captioning, input an image and proposed caption to the model, and aggregate either algebraic or learned token confidences over words or sequences to estimate image-caption consistency. Compared to sequence-level scores from pretrained models, TLC with algebraic confidence measures achieves a relative improvement in accuracy by 10% on verb understanding in SVO-Probes and outperforms prior state-of-the-art in image and group scores for compositional reasoning in Winoground by a relative 37% and 9%, respectively. When training data are available, a learned confidence estimator provides further improved performance, reducing object hallucination rates in MS COCO Captions by a relative 30% over the original model and setting a new state-of-the-art.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Reasoning | Winoground | OFA large (ITM) | Group Score | 7.25 | #66 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (ITM) | Image Score | 10.25 | #66 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (ITM) | Text Score | 30.75 | #66 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (TLC-A) | Group Score | 17.50 | #73 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (TLC-A) | Image Score | 27.00 | #73 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (TLC-A) | Text Score | 29.25 | #73 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA base (ITM) | Group Score | 6.50 | #82 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA base (ITM) | Image Score | 10.75 | #82 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA base (ITM) | Text Score | 26.75 | #82 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA base (TLC-A) | Group Score | 13.75 | #88 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA base (TLC-A) | Image Score | 23.50 | #88 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA base (TLC-A) | Text Score | 24.50 | #88 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA tiny (ITM) | Group Score | 4.50 | #94 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA tiny (ITM) | Image Score | 7.75 | #94 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA tiny (ITM) | Text Score | 22.75 | #94 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA tiny (TLC-A) | Group Score | 6.75 | #109 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA tiny (TLC-A) | Image Score | 15.75 | #109 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA tiny (TLC-A) | Text Score | 16.50 | #109 of 114 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections