{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semstyle-learning-to-generate-stylised-image","title":"SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text","arxiv_id":"1805.07030","date":"2018-05-18","proceeding":"CVPR 2018 6","authors":["Alexander Mathews","Lexing Xie","Xuming He"],"abstract":"Linguistic style is an essential part of written communication, with the\npower to affect both clarity and attractiveness. With recent advances in vision\nand language, we can start to tackle the problem of generating image captions\nthat are both visually grounded and appropriately styled. Existing approaches\neither require styled training captions aligned to images or generate captions\nwith low relevance. We develop a model that learns to generate visually\nrelevant styled captions from a large corpus of styled text without aligned\nimages. The core idea of this model, called SemStyle, is to separate semantics\nand style. One key component is a novel and concise semantic term\nrepresentation generated using natural language processing techniques and frame\nsemantics. In addition, we develop a unified language model that decodes\nsentences with diverse word choices and syntax for different styles.\nEvaluations, both automatic and manual, show captions from SemStyle preserve\nimage semantics, are descriptive, and are style shifted. More broadly, this\nwork provides possibilities to learn richer image descriptions from the\nplethora of linguistic data available on the web.","url_abs":"http://arxiv.org/abs/1805.07030v1","url_pdf":"http://arxiv.org/pdf/1805.07030v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"semstyle-learning-to-generate-stylised-image","repo_url":"https://github.com/computationalmedia/semstyle","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"descriptive","task_name":"Descriptive"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.07030","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}