{"url":"/dataset/mango","name":"mango","full_name":null,"description_markdown":"Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks. \r\nIn this paper, we propose MANGO, a benchmark to evaluate their ability to perform text-based mapping and navigation.\r\nOur benchmark includes $53$ mazes taken from a suite of textgames: each maze is paired with a walkthrough that visits every location but does **not** cover all possible paths. \r\nThe task is question-answering: for each maze, a large language model reads the walkthrough and answers hundreds of mapping and navigation questions such as \"How should you go to `Attic` from `West of House`?\" and \"Where are we if we go `north` and `east` from `Cellar`?\".\r\nAlthough these questions are easy for humans, it turns out that even GPT-4, the best-to-date language model, performs poorly when answering them. \r\nFurther, our experiments suggest that a strong mapping and navigation ability would benefit the performance of large language models on relevant downstream tasks, such as playing textgames.\r\nOur MANGO benchmark will facilitate future research on methods that improve the mapping and navigation capabilities of LLMs. \r\nWe host our leaderboard, data, code, and evaluation program at [https://mango.ttic.edu](https://mango.ttic.edu) and [https://github.com/Oaklight/mango](https://github.com/Oaklight/mango).","description_withheld":null,"homepage":"https://mango.ttic.edu","introduced_date":"2024-03-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/mango-a-benchmark-for-evaluating-mapping-and","title":"MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models","first_author":"Peng Ding","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["mango"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}