{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generating-image-descriptions-via-sequential","title":"Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze","arxiv_id":"2011.04592","date":"2020-11-09","proceeding":"EMNLP 2020 11","authors":["Ece Takmaz","Sandro Pezzelle","Lisa Beinborn","Raquel Fernández"],"abstract":"When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential cross-modal alignment by modelling the image description generation process computationally. We take as our starting point a state-of-the-art image captioning system and develop several model variants that exploit information from human gaze patterns recorded during language production. In particular, we propose the first approach to image description generation where visual processing is modelled $\\textit{sequentially}$. Our experiments and analyses confirm that better descriptions can be obtained by exploiting gaze-driven attention and shed light on human cognitive processes by comparing different ways of aligning the gaze modality with language production. We find that processing gaze data sequentially leads to descriptions that are better aligned to those produced by speakers, more diverse, and more natural${-}$particularly when gaze is encoded with a dedicated recurrent component.","url_abs":"https://arxiv.org/abs/2011.04592v1","url_pdf":"https://arxiv.org/pdf/2011.04592v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generating-image-descriptions-via-sequential","repo_url":"https://github.com/dmg-illc/didec-seq-gen","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":null,"task_name":"Image Description"},{"task_slug":"cross-modal-alignment","task_name":"cross-modal alignment"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2011.04592","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2011.04592"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/dmg-illc/didec-seq-gen","reach":null}],"summary":{"ran":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"685139212826c9d8","entry":"Attention","repo":"dmg-illc/didec-seq-gen","repo_kind":"official","path":"description_generation/models/model_pret_gaze_2RNN.py","file_url":"https://github.com/dmg-illc/didec-seq-gen/blob/HEAD/description_generation/models/model_pret_gaze_2RNN.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"685139212826c9d8"}},{"code_sha256_prefix":"dff83e37a23a8988","entry":"DecoderWithAttention","repo":"dmg-illc/didec-seq-gen","repo_kind":"official","path":"description_generation/models/model_pret_gaze_2RNN.py","file_url":"https://github.com/dmg-illc/didec-seq-gen/blob/HEAD/description_generation/models/model_pret_gaze_2RNN.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dff83e37a23a8988"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}