{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dual-cnn-a-convolutional-language-decoder-for","title":"Dual-CNN: A Convolutional language decoder for paragraph image captioning","arxiv_id":null,"date":"2020-02-14","proceeding":"Neurocomputing 2020 2","authors":["Ruifan Li","Haoyun Liang","Yihui Shi","Fangxiang Feng","Xiaojie Wang"],"abstract":"Abstract The task of paragraph image captioning aims to generate a coherent paragraph describing a given image. However, due to their limited ability to capture long-term dependency, recurrent neural network or long-short term memory based decoders could hardly generate satisfactory textual descriptions with a long paragraph. In addition, the training inefficiency in the sequential decoders is significantly observed. Motivated by the advantage of convolutional neural network (i.e., CNN), in this paper, we propose a Dual-CNN decoder with long-term memory ability and parallel computation, which can produce a semantically coherent paragraph for an image. Our Dual-CNN model is evaluated on the Stanford image-paragraph dataset. Extensive experiments demonstrate that our Dual-CNN achieves comparable results compared with state-of-the-art models. Furthermore, the diversity and coherence of generated paragraphs are analyzed to show the superiority of our approach.","url_abs":"https://www.sciencedirect.com/science/article/pii/S0925231220302319","url_pdf":"https://www.sciencedirect.com/science/article/pii/S0925231220302319","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"image-paragraph-captioning","task_name":"Image Paragraph Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-paragraph-captioning-on-image-paragraph","task":"Image Paragraph Captioning","dataset":"Image Paragraph Captioning","model":"Dual-CNN","rank_in_archive_order":8,"of":10,"metrics":{"BLEU-4":"8.6","CIDEr":"17.4","METEOR":"15.8"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}