{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hierarchical-scene-graph-encoder-decoder-for","title":"Hierarchical Scene Graph Encoder-Decoder for Image Paragraph Captioning","arxiv_id":null,"date":"2020-10-12","proceeding":"ACM International Conference on Multimedia 2020 10","authors":["Yang","Xu","Chongyang Gao","Hanwang Zhang","and Jianfei Cai"],"abstract":"When we humans tell a long paragraph about an image, we usually\r\nfirst implicitly compose a mental “script” and then comply with it\r\nto generate the paragraph. Inspired by this, we render the modern\r\nencoder-decoder based image paragraph captioning model such\r\nability by proposing Hierarchical Scene Graph Encoder-Decoder\r\n(HSGED) for generating coherent and distinctive paragraphs. In\r\nparticular, we use the image scene graph as the “script” to incorporate rich semantic knowledge and, more importantly, the hierarchical constraints into the model. Specifically, we design a sentence\r\nscene graph RNN (SSG-RNN) to generate sub-graph level topics,\r\nwhich constrain the word scene graph RNN (WSG-RNN) to generate the corresponding sentences. We propose irredundant attention\r\nin SSG-RNN to improve the possibility of abstracting topics from\r\nrarely described sub-graphs and inheriting attention in WSG-RNN\r\nto generate more grounded sentences with the abstracted topics,\r\nboth of which give rise to more distinctive paragraphs. An efficient\r\nsentence-level loss is also proposed for encouraging the sequence of\r\ngenerated sentences to be similar to that of the ground-truth paragraphs. We validate HSGED on Stanford image paragraph dataset\r\nand show that it not only achieves a new state-of-the-art 36.02\r\nCIDEr-D, but also generates more coherent and distinctive paragraphs under various metrics.","url_abs":"https://dl.acm.org/doi/abs/10.1145/3394171.3413859","url_pdf":"https://dl.acm.org/doi/pdf/10.1145/3394171.3413859","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-paragraph-captioning","task_name":"Image Paragraph Captioning"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-paragraph-captioning-on-image-paragraph","task":"Image Paragraph Captioning","dataset":"Image Paragraph Captioning","model":"HSGED(SLL)","rank_in_archive_order":1,"of":10,"metrics":{"BLEU-4":"11.26","CIDEr":"36.02","METEOR":"18.33"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}