{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/face-cap-image-captioning-using-facial","title":"Face-Cap: Image Captioning using Facial Expression Analysis","arxiv_id":"1807.02250","date":"2018-07-06","proceeding":null,"authors":["Omid Mohamad Nezami","Mark Dras","Peter Anderson","Len Hamey"],"abstract":"Image captioning is the process of generating a natural language description\nof an image. Most current image captioning models, however, do not take into\naccount the emotional aspect of an image, which is very relevant to activities\nand interpersonal relationships represented therein. Towards developing a model\nthat can produce human-like captions incorporating these, we use facial\nexpression features extracted from images including human faces, with the aim\nof improving the descriptive ability of the model. In this work, we present two\nvariants of our Face-Cap model, which embed facial expression features in\ndifferent ways, to generate image captions. Using all standard evaluation\nmetrics, our Face-Cap models outperform a state-of-the-art baseline model for\ngenerating image captions when applied to an image caption dataset extracted\nfrom the standard Flickr 30K dataset, consisting of around 11K images\ncontaining faces. An analysis of the captions finds that, perhaps surprisingly,\nthe improvement in caption quality appears to come not from the addition of\nadjectives linked to emotional aspects of the images, but from more variety in\nthe actions described in the captions.","url_abs":"http://arxiv.org/abs/1807.02250v2","url_pdf":"http://arxiv.org/pdf/1807.02250v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"face-cap-image-captioning-using-facial","repo_url":"https://github.com/omidmn/Face-Cap","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"descriptive","task_name":"Descriptive"},{"task_slug":"image-captioning","task_name":"Image Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}