{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-lambada-dataset-word-prediction-requiring","title":"The LAMBADA dataset: Word prediction requiring a broad discourse context","arxiv_id":"1606.06031","date":"2016-06-20","proceeding":"ACL 2016 8","authors":["Denis Paperno","Germán Kruszewski","Angeliki Lazaridou","Quan Ngoc Pham","Raffaella Bernardi","Sandro Pezzelle","Marco Baroni","Gemma Boleda","Raquel Fernández"],"abstract":"We introduce LAMBADA, a dataset to evaluate the capabilities of computational\nmodels for text understanding by means of a word prediction task. LAMBADA is a\ncollection of narrative passages sharing the characteristic that human subjects\nare able to guess their last word if they are exposed to the whole passage, but\nnot if they only see the last sentence preceding the target word. To succeed on\nLAMBADA, computational models cannot simply rely on local context, but must be\nable to keep track of information in the broader discourse. We show that\nLAMBADA exemplifies a wide range of linguistic phenomena, and that none of\nseveral state-of-the-art language models reaches accuracy above 1% on this\nnovel benchmark. We thus propose LAMBADA as a challenging test set, meant to\nencourage the development of new models capable of genuine understanding of\nbroad context in natural language text.","url_abs":"http://arxiv.org/abs/1606.06031v1","url_pdf":"http://arxiv.org/pdf/1606.06031v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-lambada-dataset-word-prediction-requiring","repo_url":"https://github.com/jdeschena/sdtt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"the-lambada-dataset-word-prediction-requiring","repo_url":"https://github.com/keyonvafa/sequential-rationales","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":null},{"paper_slug":"the-lambada-dataset-word-prediction-requiring","repo_url":"https://github.com/zhenwang9102/coherence-boosting","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"lambada","task_name":"LAMBADA"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[{"slug":"lambada","name":"LAMBADA","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1606.06031","atlas_url":"https://app.syntology.ai/?focus=1606.06031","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}