{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scene-graph-generation-from-objects-phrases","title":"Scene Graph Generation from Objects, Phrases and Region Captions","arxiv_id":"1707.09700","date":"2017-07-31","proceeding":"ICCV 2017 10","authors":["Yikang Li","Wanli Ouyang","Bolei Zhou","Kun Wang","Xiaogang Wang"],"abstract":"Object detection, scene graph generation and region captioning, which are\nthree scene understanding tasks at different semantic levels, are tied\ntogether: scene graphs are generated on top of objects detected in an image\nwith their pairwise relationship predicted, while region captioning gives a\nlanguage description of the objects, their attributes, relations, and other\ncontext information. In this work, to leverage the mutual connections across\nsemantic levels, we propose a novel neural network model, termed as Multi-level\nScene Description Network (denoted as MSDN), to solve the three vision tasks\njointly in an end-to-end manner. Objects, phrases, and caption regions are\nfirst aligned with a dynamic graph based on their spatial and semantic\nconnections. Then a feature refining structure is used to pass messages across\nthe three levels of semantic tasks through the graph. We benchmark the learned\nmodel on three tasks, and show the joint learning across three tasks with our\nproposed method can bring mutual improvements over previous models.\nParticularly, on the scene graph generation task, our proposed method\noutperforms the state-of-art method with more than 3% margin.","url_abs":"http://arxiv.org/abs/1707.09700v2","url_pdf":"http://arxiv.org/pdf/1707.09700v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scene-graph-generation-from-objects-phrases","repo_url":"https://github.com/yikang-li/MSDN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"graph-generation","task_name":"Graph Generation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"scene-graph-generation","task_name":"Scene Graph Generation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/object-detection-on-visual-genome","task":"Object Detection","dataset":"Visual Genome","model":"MSDN","rank_in_archive_order":3,"of":4,"metrics":{"MAP":"7.43"},"uses_additional_data":false},{"leaderboard":"/sota/scene-graph-generation-on-visual-genome","task":"Scene Graph Generation","dataset":"Visual Genome","model":"MSDN","rank_in_archive_order":14,"of":19,"metrics":{"Recall@50":"10.72"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.09700","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}