Browse State-of-the-Art › Vision and Language Navigation
Vision and Language Navigation
114 papers with code · 5 benchmarks · 13 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| VLN Challenge (145 rows) | human | — | — | — | Compare |
| Touchdown Dataset (12 rows) | FLAME | FLAME: Learning to Navigate with Multimodal LLM in Urban Environments | code | Syntology ran 6 of 7 samples · 1 unverified | Compare |
| RxR (6 rows) | MARVAL | A New Path: Scaling Vision-and-Language Navigation with Synthetic... | — | — | Compare |
| map2seq (5 rows) | FLAME | FLAME: Learning to Navigate with Multimodal LLM in Urban Environments | code | Syntology ran 6 of 7 samples · 1 unverified | Compare |
| robo-vln (1 row) | Hierarchical Cross-Modal Agent | Hierarchical Cross-Modal Agent for Robotics Vision-and-Language Navigation | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
13 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 114 papers with code (223 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Nov 2017 8 repositories listedThis is significant because a robot interpreting a natural-language navigation instruction on the basis of what it sees is carrying out a vision and language process that is similar to Visual Question Answering.
-
6 Apr 2020 5 repositories listedWe develop a language-guided navigation task set in a continuous 3D environment where agents must execute low-level actions to follow natural language navigation directions.
-
13 Jul 2021 4 repositories listed Syntology ran 6 of 10 samples · 4 unverified · 9 pointer-only (licence)Most existing Vision-and-Language (V&L) models rely on pre-trained visual encoders, using a relatively small set of manually-annotated data (as compared to web-crawled data), to perceive the visual world.
-
10 Jan 2020 4 repositories listed Syntology ran 1 of 8 samples · 7 unverified · 1 pointer-only (licence)These have been added to the StreetLearn dataset and can be obtained via the same process as used previously for StreetLearn.
-
29 Nov 2018 4 repositories listedWe study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task.
-
15 Oct 2020 3 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedWe introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset.
-
5 Mar 2019 3 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedAs deep learning continues to make progress for challenging perception tasks, there is increased interest in combining vision, language, and decision-making.
-
18 Mar 2024 2 repositories listedTo address this issue, we propose a Hierarchical Spatial Proximity Reasoning (HSPR) method.
-
8 Feb 2024 2 repositories listed Syntology ran 21 of 27 samples · 6 unverifiedWe propose the problem of conversational web navigation, where a digital agent controls a web browser and follows user instructions to solve real-world tasks in a multi-turn dialogue fashion.
-
26 May 2023 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 2 pointer-only (licence)Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling.
-
20 Aug 2021 2 repositories listed Syntology ran 7 of 14 samples · 7 unverifiedGiven the scarcity of domain-specific training data and the high diversity of image and language inputs, the generalization of VLN agents to unseen environments remains challenging.
-
10 Jan 2019 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThe Vision-and-Language Navigation (VLN) task entails an agent following navigational instruction in photo-realistic unknown environments.
-
30 Jun 2025 1 repository listed Syntology ran 6 of 12 samples · 6 unverified · 12 pointer-only (licence)Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions.
-
11 Jun 2025 1 repository listedVision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, unfamiliar environments.
-
27 May 2025 1 repository listedFurthermore, we introduce a cross-interaction mechanism to regularize the imagined outputs and inject them into a navigation expert module, allowing ATD to jointly exploit both the reasoning capacity of the LLM and the…
-
19 May 2025 1 repository listedUnmanned Aerial Vehicle (UAV) Vision-and-Language Navigation (VLN) is vital for applications such as disaster response, logistics delivery, and urban inspection.
-
16 May 2025 1 repository listedVision-and-Language Navigation (VLN) is a core task where embodied agents leverage their spatial mobility to navigate in 3D environments toward designated destinations based on natural language instructions.
-
8 May 2025 1 repository listedIn this work, we propose \textbf{CityNavAgent}, a large language model (LLM)-empowered agent that significantly reduces the navigation complexity for urban aerial VLN.
-
16 Feb 2025 1 repository listedWe annotate over 2 million navigation instructions across 861 scenes and evaluate the data quality and navigation performance of trained models.
-
29 Jan 2025 1 repository listed Syntology ran 2 of 14 samples · 12 unverifiedTo evaluate the proposed task, one has to address two challenges in existing VLN datasets: the lack of OOD data, and the limited number and style diversity of instructions for each scene.
-
9 Dec 2024 1 repository listedSUSA includes a Textual Semantic Understanding (TSU) module, which narrows the modality gap between instructions and environments by generating and associating the descriptions of environmental landmarks in agent's…
-
26 Nov 2024 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)Furthermore, we prepare a large-scale 3D-language dataset to align the representations of the feature fields with language.
-
9 Sep 2024 1 repository listedEmbodied AI aims to develop robots that can \textit{understand} and execute human language instructions, as well as communicate in natural languages.
-
20 Aug 2024 1 repository listed Syntology ran 6 of 7 samples · 1 unverifiedLarge Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges.
-
19 Aug 2024 1 repository listedFirst, VLN-CE agents that discretize the visual environment are primarily trained with high-level view selection, which causes them to ignore crucial spatial reasoning within the low-level action movements.
-
31 Jul 2024 1 repository listedReal-world navigation often involves dealing with unexpected obstructions such as closed doors, moved objects, and unpredictable entities.
-
17 Jul 2024 1 repository listed Syntology ran 12 of 14 samples · 2 unverifiedCapitalizing on the remarkable advancements in Large Language Models (LLMs), there is a burgeoning initiative to harness LLMs for instruction following robotic navigation.
-
16 Jul 2024 1 repository listedIn this work, we propose an alternative method that facilitates navigation planning by considering the alignment between instructions and directed fidelity trajectories, which refers to a path from the initial node to…
-
9 Jul 2024 1 repository listedVision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development.
-
28 Jun 2024 1 repository listedHowever, performance substantially drops in new environments with no training data.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections