Datasets › CHOCOLATE

CHOCOLATE (Captions Have Often ChOsen Lies About The Evidence)

Introduced by Kung-Hsiang Huang et al. in Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning15 Dec 2023 archive 2025-07-28

CHOCOLATE is a benchmark for detecting and correcting factual inconsistency in generated chart captions. It consists of captions produced by six advanced models, which are categorized into three subsets:

  • LVLM: GPT-4V, Bard (before Gemini)
  • LLM-based Pipeline: DePlot + GPT-4
  • Fine-tuned Model: ChartT5, MatCha, UniChart

The charts are from two datasets: VisText and the Pew split of Chart-to-Text. In total, CHOCOLATE consists of 1,187 examples. Each instance in CHOCOLATE consists of a caption generated by one of the models and the annotations of the factual errors for each caption sentence.

Paper Information
Citation

If you use the CHOCOLATE dataset in your work, please kindly cite the paper using this BibTeX:

@misc{huang-etal-2023-do,
    title = "Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning",
    author = "Huang, Kung-Hsiang  and
      Zhou, Mingyang and
      Chan, Hou Pong  and
      Fung, Yi R. and
      Wang, Zhenhailong and
      Zhang, Lingyu and
      Chang, Shih-Fu and
      Ji, Heng",
    year={2023},
    eprint={2312.10160},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}    

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Factual Inconsistency Detection in Chart Captioning CHOCOLATE-LLM GPT-4V Kendall's Tau-c 0.205 GPT-4 Technical Report openai/evals +10 5 Compare
Factual Inconsistency Detection in Chart Captioning CHOCOLATE-FT Bard (before Gemini) Kendall's Tau-c 0.291 — — 5 Compare
Factual Inconsistency Detection in Chart Captioning CHOCOLATE-LVLM ChartVE Kendall's Tau-c 0.178 Do LVLMs Understand Charts? Analyzing and Correcting... huggingface/transformers +2 5 Compare
Factual Inconsistency Detection in Chart Captioning CHOCOLATE ChartVE Kendall's Tau-c 0.178 Do LVLMs Understand Charts? Analyzing and Correcting... huggingface/transformers +2 1 Compare

Papers archive 2025-07-28

4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 4. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning 3 4 15 Dec 2023 not harvested
Improved Baselines with Visual Instruction Tuning 9 3 5 Oct 2023 ran 6 of 9 samples (3 unverified; 8 pointer-only for licence)
GPT-4 Technical Report 11 1 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
DePlot: One-shot visual language reasoning by plot-to-table translation 1 3 20 Dec 2022 not harvested

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Apache-2.0 license

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CHOCOLATE
  • CHOCOLATE-LLM
  • CHOCOLATE-FT
  • CHOCOLATE-LVLM

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections