{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/goat-bench-a-benchmark-for-multi-modal","title":"GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation","arxiv_id":"2404.06609","date":"2024-04-09","proceeding":"CVPR 2024 1","authors":["Mukul Khanna","Ram Ramrakhya","Gunjan Chhablani","Sriram Yenamandra","Theophile Gervet","Matthew Chang","Zsolt Kira","Devendra Singh Chaplot","Dhruv Batra","Roozbeh Mottaghi"],"abstract":"The Embodied AI community has made significant strides in visual navigation tasks, exploring targets from 3D coordinates, objects, language descriptions, and images. However, these navigation models often handle only a single input modality as the target. With the progress achieved so far, it is time to move towards universal navigation models capable of handling various goal types, enabling more effective user interaction with robots. To facilitate this goal, we propose GOAT-Bench, a benchmark for the universal navigation task referred to as GO to AnyThing (GOAT). In this task, the agent is directed to navigate to a sequence of targets specified by the category name, language description, or image in an open-vocabulary fashion. We benchmark monolithic RL and modular methods on the GOAT task, analyzing their performance across modalities, the role of explicit and implicit scene memories, their robustness to noise in goal specifications, and the impact of memory in lifelong scenarios.","url_abs":"https://arxiv.org/abs/2404.06609v1","url_pdf":"https://arxiv.org/pdf/2404.06609v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"goat-bench-a-benchmark-for-multi-modal","repo_url":"https://github.com/Ram81/goat-bench","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"go-to-anything","task_name":"Go to AnyThing"},{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"universal-navigation","task_name":"Universal Navigation"},{"task_slug":"visual-navigation","task_name":"Visual Navigation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2404.06609","atlas_url":"https://app.syntology.ai/?focus=2404.06609","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2404.06609"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Ram81/goat-bench","reach":{"status":"ok"}}],"summary":{"ran":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"b433d0fa97ed9b6a","entry":"forward_avg_attn_pool","repo":"Ram81/goat-bench","repo_kind":"official","path":"goat_bench/models/clip_policy.py","file_url":"https://github.com/Ram81/goat-bench/blob/HEAD/goat_bench/models/clip_policy.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b433d0fa97ed9b6a"}},{"code_sha256_prefix":"56e951aecf753858","entry":"get_transform","repo":"Ram81/goat-bench","repo_kind":"official","path":"goat_bench/models/transforms.py","file_url":"https://github.com/Ram81/goat-bench/blob/HEAD/goat_bench/models/transforms.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"56e951aecf753858"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}