{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-the-synergy-between-vision-language","title":"Exploring the Synergy Between Vision-Language Pretraining and ChatGPT for Artwork Captioning: A Preliminary Study","arxiv_id":null,"date":"2023-01-21","proceeding":"FAPER Workshop - ICIAP 2023 2023 1","authors":["Giovanna Castellano","Nicola Fanelli","Raffaele Scaringi","Gennaro Vessio"],"abstract":"While AI techniques have enabled automated analysis and interpretation of visual content, generating meaningful captions for artworks presents unique challenges. These include understanding artistic intent, historical context, and complex visual elements. Despite recent developments in multi-modal techniques, there are still gaps in generating complete and accurate captions. This paper contributes by introducing a new dataset for artwork captioning generated using prompt engineering techniques and ChatGPT. We refined the captions with CLIPScore to filter out noise; then, we fine-tuned GIT-Base, resulting in visually accurate captions that surpass the ground truth. Enrichment of descriptions with predicted metadata improves their informativeness. Artwork captioning has implications for art appreciation, inclusivity, education, and cultural exchange, particularly for people with visual impairments or limited knowledge of art.","url_abs":"https://www.researchgate.net/publication/377558832_Exploring_the_Synergy_Between_Vision-Language_Pretraining_and_ChatGPT_for_Artwork_Captioning_A_Preliminary_Study","url_pdf":"https://www.researchgate.net/profile/Gennaro-Vessio/publication/377558832_Exploring_the_Synergy_Between_Vision-Language_Pretraining_and_ChatGPT_for_Artwork_Captioning_A_Preliminary_Study/links/65acc767bf5b00662e2ffc4a/Exploring-the-Synergy-Between-Vision-Language-Pretraining-and-ChatGPT-for-Artwork-Captioning-A-Preliminary-Study.pdf?origin=publicationDetail&_sg%5B0%5D=W5KnRbp6cLdH8O_QzDIdxs35wF-7NiBkmGVuyaGUH9okj5sk8-xta8DS5zjLb808KM4sVb6NUqPtjZyyOtiHeg.MTxS0S6ohb0tA4_E7O0bKcszJl2B3H3SX_4-Hf9NEKChg-L-L4AdK9UId7o7-cVrVefYPf8_8qze9H-5unOjUw&_sg%5B1%5D=AsBuFzQff6HuD1dQNUmCNQRJ21CJgSeKuUMAVxwutbed6XvZlnfJcrATVpN74lLTMCKTiRpKqzVrIYANwQ1D0rRV9HO_BOLFlfMBe5tct6w-.MTxS0S6ohb0tA4_E7O0bKcszJl2B3H3SX_4-Hf9NEKChg-L-L4AdK9UId7o7-cVrVefYPf8_8qze9H-5unOjUw&_iepl=&_rtd=eyJjb250ZW50SW50ZW50IjoibWFpbkl0ZW0ifQ%3D%3D&_tp=eyJjb250ZXh0Ijp7ImZpcnN0UGFnZSI6InB1YmxpY2F0aW9uIiwicGFnZSI6InB1YmxpY2F0aW9uIiwicG9zaXRpb24iOiJwYWdlSGVhZGVyIn19","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exploring-the-synergy-between-vision-language","repo_url":"https://github.com/nicolafan/neural-artwork-caption-generator","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"informativeness","task_name":"Informativeness"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"prompt-engineering","task_name":"Prompt Engineering"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}