{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/neural-self-talk-image-understanding-via","title":"Neural Self Talk: Image Understanding via Continuous Questioning and Answering","arxiv_id":"1512.03460","date":"2015-12-10","proceeding":null,"authors":["Yezhou Yang","Yi Li","Cornelia Fermuller","Yiannis Aloimonos"],"abstract":"In this paper we consider the problem of continuously discovering image\ncontents by actively asking image based questions and subsequently answering\nthe questions being asked. The key components include a Visual Question\nGeneration (VQG) module and a Visual Question Answering module, in which\nRecurrent Neural Networks (RNN) and Convolutional Neural Network (CNN) are\nused. Given a dataset that contains images, questions and their answers, both\nmodules are trained at the same time, with the difference being VQG uses the\nimages as input and the corresponding questions as output, while VQA uses\nimages and questions as input and the corresponding answers as output. We\nevaluate the self talk process subjectively using Amazon Mechanical Turk, which\nshow effectiveness of the proposed method.","url_abs":"http://arxiv.org/abs/1512.03460v1","url_pdf":"http://arxiv.org/pdf/1512.03460v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"question-generation","task_name":"Question Generation"},{"task_slug":"question-generation","task_name":"Question-Generation"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/question-generation-on-coco-visual-question","task":"Question Generation","dataset":"COCO Visual Question Answering (VQA) real images 1.0 open ended","model":"Max(Yang,2015)","rank_in_archive_order":3,"of":4,"metrics":{"BLEU-1":"59.4"},"uses_additional_data":false},{"leaderboard":"/sota/question-generation-on-coco-visual-question","task":"Question Generation","dataset":"COCO Visual Question Answering (VQA) real images 1.0 open ended","model":"Sample(Yang,2015)","rank_in_archive_order":4,"of":4,"metrics":{"BLEU-1":"38.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1512.03460","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}