{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-compositional-captioning-describing","title":"Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data","arxiv_id":"1511.05284","date":"2015-11-17","proceeding":"CVPR 2016 6","authors":["Lisa Anne Hendricks","Subhashini Venugopalan","Marcus Rohrbach","Raymond Mooney","Kate Saenko","Trevor Darrell"],"abstract":"While recent deep neural network models have achieved promising results on\nthe image captioning task, they rely largely on the availability of corpora\nwith paired image and sentence captions to describe objects in context. In this\nwork, we propose the Deep Compositional Captioner (DCC) to address the task of\ngenerating descriptions of novel objects which are not present in paired\nimage-sentence datasets. Our method achieves this by leveraging large object\nrecognition datasets and external text corpora and by transferring knowledge\nbetween semantically similar concepts. Current deep caption models can only\ndescribe objects contained in paired image-sentence corpora, despite the fact\nthat they are pre-trained with large object recognition datasets, namely\nImageNet. In contrast, our model can compose sentences that describe novel\nobjects and their interactions with other objects. We demonstrate our model's\nability to describe novel concepts by empirically evaluating its performance on\nMSCOCO and show qualitative results on ImageNet images of objects for which no\npaired image-caption data exist. Further, we extend our approach to generate\ndescriptions of objects in video clips. Our results show that DCC has distinct\nadvantages over existing image and video captioning approaches for generating\ndescriptions of new objects in context.","url_abs":"http://arxiv.org/abs/1511.05284v2","url_pdf":"http://arxiv.org/pdf/1511.05284v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-compositional-captioning-describing","repo_url":"https://github.com/LisaAnne/DCC","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"novel-concepts","task_name":"Novel Concepts"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"video-captioning","task_name":"Video Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.05284","atlas_url":"https://app.syntology.ai/?focus=1511.05284","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}