{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-role-of-chain-of-thought-in-complex","title":"The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task","arxiv_id":"2311.09193","date":"2023-11-15","proceeding":null,"authors":["Yifan Wu","Pengchuan Zhang","Wenhan Xiong","Barlas Oguz","James C. Gee","Yixin Nie"],"abstract":"The study explores the effectiveness of the Chain-of-Thought approach, known for its proficiency in language tasks by breaking them down into sub-tasks and intermediate steps, in improving vision-language tasks that demand sophisticated perception and reasoning. We present the \"Description then Decision\" strategy, which is inspired by how humans process signals. This strategy significantly improves probing task performance by 50%, establishing the groundwork for future research on reasoning paradigms in complex vision-language tasks.","url_abs":"https://arxiv.org/abs/2311.09193v1","url_pdf":"https://arxiv.org/pdf/2311.09193v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"GPT-4V (CoT, pick b/w two options)","rank_in_archive_order":2,"of":114,"metrics":{"Group Score":"58.75","Image Score":"68.75","Text Score":"75.25"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"GPT-4V (pick b/w two options)","rank_in_archive_order":3,"of":114,"metrics":{"Group Score":"39.25","Image Score":"46.25","Text Score":"69.25"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2311.09193","atlas_url":"https://app.syntology.ai/?focus=2311.09193","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}