{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deplot-one-shot-visual-language-reasoning-by","title":"DePlot: One-shot visual language reasoning by plot-to-table translation","arxiv_id":"2212.10505","date":"2022-12-20","proceeding":null,"authors":["Fangyu Liu","Julian Martin Eisenschlos","Francesco Piccinno","Syrine Krichene","Chenxi Pang","Kenton Lee","Mandar Joshi","Wenhu Chen","Nigel Collier","Yasemin Altun"],"abstract":"Visual language such as charts and plots is ubiquitous in the human world. Comprehending plots and charts requires strong reasoning skills. Prior state-of-the-art (SOTA) models require at least tens of thousands of training examples and their reasoning capabilities are still much limited, especially on complex human-written queries. This paper presents the first one-shot solution to visual language reasoning. We decompose the challenge of visual language reasoning into two steps: (1) plot-to-text translation, and (2) reasoning over the translated text. The key in this method is a modality conversion module, named as DePlot, which translates the image of a plot or chart to a linearized table. The output of DePlot can then be directly used to prompt a pretrained large language model (LLM), exploiting the few-shot reasoning capabilities of LLMs. To obtain DePlot, we standardize the plot-to-table task by establishing unified task formats and metrics, and train DePlot end-to-end on this task. DePlot can then be used off-the-shelf together with LLMs in a plug-and-play fashion. Compared with a SOTA model finetuned on more than >28k data points, DePlot+LLM with just one-shot prompting achieves a 24.0% improvement over finetuned SOTA on human-written queries from the task of chart QA.","url_abs":"https://arxiv.org/abs/2212.10505v2","url_pdf":"https://arxiv.org/pdf/2212.10505v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deplot-one-shot-visual-language-reasoning-by","repo_url":"https://github.com/huggingface/transformers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"chart-question-answering","task_name":"Chart Question Answering"},{"task_slug":"factual-inconsistency-detection-in-chart","task_name":"Factual Inconsistency Detection in Chart Captioning"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/chart-question-answering-on-chartqa","task":"Chart Question Answering","dataset":"ChartQA","model":"DePlot+FlanPaLM+Codex (PoT Self-Consistency)","rank_in_archive_order":3,"of":27,"metrics":{"1:1 Accuracy":"79.3"},"uses_additional_data":false},{"leaderboard":"/sota/chart-question-answering-on-chartqa","task":"Chart Question Answering","dataset":"ChartQA","model":"DePlot+Codex (PoT Self-Consistency)","rank_in_archive_order":5,"of":27,"metrics":{"1:1 Accuracy":"76.7"},"uses_additional_data":false},{"leaderboard":"/sota/chart-question-answering-on-chartqa","task":"Chart Question Answering","dataset":"ChartQA","model":"DePlot+FlanPaLM (Self-Consistency)","rank_in_archive_order":13,"of":27,"metrics":{"1:1 Accuracy":"70.5"},"uses_additional_data":false},{"leaderboard":"/sota/chart-question-answering-on-chartqa","task":"Chart Question Answering","dataset":"ChartQA","model":"DePlot+FlanPaLM (CoT)","rank_in_archive_order":16,"of":27,"metrics":{"1:1 Accuracy":"67.3"},"uses_additional_data":false},{"leaderboard":"/sota/chart-question-answering-on-chartqa","task":"Chart Question Answering","dataset":"ChartQA","model":"DePlot+GPT3 (Self-Consistency)","rank_in_archive_order":26,"of":27,"metrics":{"1:1 Accuracy":"42.3"},"uses_additional_data":false},{"leaderboard":"/sota/chart-question-answering-on-chartqa","task":"Chart Question Answering","dataset":"ChartQA","model":"DePlot+GPT3 (CoT)","rank_in_archive_order":27,"of":27,"metrics":{"1:1 Accuracy":"36.9"},"uses_additional_data":false},{"leaderboard":"/sota/chart-question-answering-on-plotqa","task":"Chart Question Answering","dataset":"PlotQA","model":"DePlot+FlanPaLM+Codex\n(PoT Self-Consistency)","rank_in_archive_order":3,"of":6,"metrics":{"1:1 Accuracy":"66.6"},"uses_additional_data":false},{"leaderboard":"/sota/factual-inconsistency-detection-in-chart-2","task":"Factual Inconsistency Detection in Chart Captioning","dataset":"CHOCOLATE-FT","model":"DePlot + GPT-4","rank_in_archive_order":5,"of":5,"metrics":{"Kendall's Tau-c":"0.109"},"uses_additional_data":false},{"leaderboard":"/sota/factual-inconsistency-detection-in-chart-1","task":"Factual Inconsistency Detection in Chart Captioning","dataset":"CHOCOLATE-LLM","model":"DePlot + GPT-4","rank_in_archive_order":2,"of":5,"metrics":{"Kendall's Tau-c":"0.117"},"uses_additional_data":false},{"leaderboard":"/sota/factual-inconsistency-detection-in-chart-3","task":"Factual Inconsistency Detection in Chart Captioning","dataset":"CHOCOLATE-LVLM","model":"DePlot + GPT-4","rank_in_archive_order":3,"of":5,"metrics":{"Kendall's Tau-c":"0.129"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2212.10505","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}