Papers › Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning

Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning

15 Dec 2023arXiv:2312.10160archive 2025-07-28

Kung-Hsiang Huang, Mingyang Zhou, Hou Pong Chan, Yi R. Fung, Zhenhailong Wang, Lingyu Zhang, Shih-Fu Chang, Heng Ji

Recent advancements in large vision-language models (LVLMs) have led to significant progress in generating natural language descriptions for visual content and thus enhancing various applications. One issue with these powerful models is that they sometimes produce texts that are factually inconsistent with the visual input. While there has been some effort to mitigate such inconsistencies in natural image captioning, the factuality of generated captions for structured document images, such as charts, has not received as much scrutiny, posing a potential threat to information reliability in critical applications. This work delves into the factuality aspect by introducing a comprehensive typology of factual errors in generated chart captions. A large-scale human annotation effort provides insight into the error patterns and frequencies in captions crafted by various chart captioning models, ultimately forming the foundation of a novel dataset, CHOCOLATE. Our analysis reveals that even state-of-the-art models, including GPT-4V, frequently produce captions laced with factual inaccuracies. In response to this challenge, we establish the new task of Chart Caption Factual Error Correction and introduce CHARTVE, a model for visual entailment that outperforms proprietary and open-source LVLMs in evaluating factual consistency. Furthermore, we propose C2TFEC, an interpretable two-stage framework that excels at correcting factual errors. This work inaugurates a new domain in factual error correction for chart captions, presenting a novel evaluation mechanism, and demonstrating an effective approach to ensuring the factuality of generated chart captions. The code and data as well as the continuously updated benchmark can be found at: https://khuangaf.github.io/CHOCOLATE/.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

khuangaf/chocolate officialmentioned in papermentioned on GitHubpytorch report
salesforceairesearch/crmarena mentioned on GitHubNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Factual Inconsistency Detection in Chart CaptioningImage CaptioningVisual Entailment

Datasets

Introduced by this paper, per the archive.

CHOCOLATE

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Factual Inconsistency Detection in Chart Captioning CHOCOLATE ChartVE Kendall's Tau-c 0.178 #1 of 1 Archive leaderboard report
Factual Inconsistency Detection in Chart Captioning CHOCOLATE-FT ChartVE Kendall's Tau-c 0.215 #2 of 5 Archive leaderboard report
Factual Inconsistency Detection in Chart Captioning CHOCOLATE-LLM ChartVE Kendall's Tau-c 0.091 #4 of 5 Archive leaderboard report
Factual Inconsistency Detection in Chart Captioning CHOCOLATE-LVLM ChartVE Kendall's Tau-c 0.178 #1 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections