{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-fill-in-the-blank-using-lrrl-lstms-with","title":"Video Fill In the Blank using LR/RL LSTMs with Spatial-Temporal Attentions","arxiv_id":"1704.04689","date":"2017-04-15","proceeding":"ICCV 2017 10","authors":["Amir Mazaheri","Dong Zhang","Mubarak Shah"],"abstract":"Given a video and a description sentence with one missing word (we call it\nthe \"source sentence\"), Video-Fill-In-the-Blank (VFIB) problem is to find the\nmissing word automatically. The contextual information of the sentence, as well\nas visual cues from the video, are important to infer the missing word\naccurately. Since the source sentence is broken into two fragments: the\nsentence's left fragment (before the blank) and the sentence's right fragment\n(after the blank), traditional Recurrent Neural Networks cannot encode this\nstructure accurately because of many possible variations of the missing word in\nterms of the location and type of the word in the source sentence. For example,\na missing word can be the first word or be in the middle of the sentence and it\ncan be a verb or an adjective. In this paper, we propose a framework to tackle\nthe textual encoding: Two separate LSTMs (the LR and RL LSTMs) are employed to\nencode the left and right sentence fragments and a novel structure is\nintroduced to combine each fragment with an \"external memory\" corresponding the\nopposite fragments. For the visual encoding, end-to-end spatial and temporal\nattention models are employed to select discriminative visual representations\nto find the missing word. In the experiments, we demonstrate the superior\nperformance of the proposed method on challenging VFIB problem. Furthermore, we\nintroduce an extended and more generalized version of VFIB, which is not\nlimited to a single blank. Our experiments indicate the generalization\ncapability of our method in dealing with such more realistic scenarios.","url_abs":"http://arxiv.org/abs/1704.04689v1","url_pdf":"http://arxiv.org/pdf/1704.04689v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-fill-in-the-blank-using-lrrl-lstms-with","repo_url":"https://github.com/amirmazaheri1990/VFIB-LRRLLSTMs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1704.04689","atlas_url":"https://app.syntology.ai/?focus=1704.04689","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}