Browse State-of-the-Art › Vision-Language Navigation
Vision-Language Navigation
38 papers with code · 1 benchmark · 9 datasets archive 2025-07-28
Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments.
( Image credit: Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Room2Room (3 rows) | R2R+EnvDrop | Learning to Navigate Unseen Environments: Back Translation with... | code | Syntology ran 0 of 4 samples · 4 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
9 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 38 papers with code (81 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Mar 2019 3 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedAs deep learning continues to make progress for challenging perception tasks, there is increased interest in combining vision, language, and decision-making.
-
24 Oct 2019 2 repositories listedCommanding a robot to navigate with natural language instructions is a long-term goal for grounded language understanding and robotics.
-
10 Jan 2019 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThe Vision-and-Language Navigation (VLN) task entails an agent following navigational instruction in photo-realistic unknown environments.
-
20 Jun 2025 1 repository listedVision-Language Navigation (VLN) is a core challenge in embodied AI, requiring agents to navigate real-world environments using natural language instructions.
-
23 Mar 2025 1 repository listedExperiments on both the discrete environments (R2R, REVERIE, and R4R datasets) and continuous environments (R2R-CE dataset) show the superior performance and impressive generalization ability of our method.
-
9 Dec 2024 1 repository listedSUSA includes a Textual Semantic Understanding (TSU) module, which narrows the modality gap between instructions and environments by generating and associating the descriptions of environmental landmarks in agent's…
-
8 Dec 2024 1 repository listedHuman-interactive robotic systems, particularly autonomous vehicles (AVs), must effectively integrate human instructions into their motion planning.
-
DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control20 Jul 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedBuilding a general-purpose intelligent home-assistant agent skilled in diverse tasks by human commands is a long-term blueprint of embodied AI research, which poses requirements on task planning, environment modeling,…
-
29 May 2024 1 repository listedTo mitigate the noise in the priors due to the lack of visual constraints, we introduce a learnable cooccurrence scoring module, which corrects the importance of each cooccurrence according to actual observations for…
-
27 Apr 2024 1 repository listedHumans excel at forming mental maps of their surroundings, equipping them to understand object relationships and navigate based on language queries.
-
2 Apr 2024 1 repository listed Syntology ran 6 of 12 samples · 6 unverified · 12 pointer-only (licence)Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments.
-
21 Mar 2024 1 repository listedTo achieve a comprehensive 3D representation with fine-grained details, we introduce a Volumetric Environment Representation (VER), which voxelizes the physical world into structured 3D cells.
-
2 Dec 2023 1 repository listed Syntology ran 6 of 11 samples · 5 unverified · 11 pointer-only (licence)In this paper, we aim to tackle this problem with a unified framework consisting of an end-to-end trainable method and a planning algorithm.
-
18 Nov 2023 1 repository listed Syntology ran 3 of 13 samples · 10 unverifiedHowever, several significant challenges remain: (i) most of these models rely on 2D images yet exhibit a limited capacity for 3D input; (ii) these models rarely explore the tasks inherently defined in 3D world, e.
-
9 Aug 2023 1 repository listed Syntology ran 10 of 14 samples · 4 unverified · 14 pointer-only (licence)Vision-language navigation (VLN), which entails an agent to navigate 3D environments following human instructions, has shown great advances.
-
6 Apr 2023 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedTo develop a robust VLN-CE agent, we propose a new navigation framework, ETPNav, which focuses on two critical skills: 1) the capability to abstract environments and generate long-range navigation plans, and 2) the…
-
1 Jan 2023 1 repository listedIn this paper, we propose an Adaptive Zone-aware Hierarchical Planner (AZHP) to explicitly divides the navigation process into two heterogeneous phases, i.
-
30 Oct 2022 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedWith the emergence of varied visual navigation tasks (e.
-
22 Oct 2022 1 repository listed Syntology ran 1 of 6 samples · 5 unverifiedThese reactive agents are insufficient for long-horizon complex tasks.
-
19 Jul 2022 1 repository listedVision-language navigation is the task of directing an embodied agent to navigate in 3D scenes with natural language instructions.
-
CLEAR: Improving Vision-Language Navigation with Cross-Lingual, Environment-Agnostic Representations5 Jul 2022 1 repository listedEmpirically, on the Room-Across-Room dataset, we show that our multilingual agent gets large improvements in all metrics over the strong baseline model when generalizing to unseen environments with the cross-lingual…
-
20 Apr 2022 1 repository listedHowever, the crucial navigation clues (i.
-
30 Mar 2022 1 repository listedSince the rise of vision-language navigation (VLN), great progress has been made in instruction following -- building a follower to navigate environments under the guidance of instructions.
-
8 Mar 2022 1 repository listedTo improve the ability of fast cross-domain adaptation, we propose Prompt-based Environmental Self-exploration (ProbES), which can self-explore the environments by sampling trajectories and automatically generates…
-
4 Feb 2022 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)To study VLN with unknown command feasibility, we introduce a new dataset Mobile app Tasks with Iterative Feedback (MoTIF), where the goal is to complete a natural language command in a mobile app.
-
8 Dec 2021 1 repository listedThe vision-language navigation (VLN) task requires an agent to reach a target with the guidance of natural language instruction.
-
23 Jul 2021 1 repository listedSpecifically, we propose a Dynamic Reinforced Instruction Attacker (DR-Attacker), which learns to mislead the navigator to move to the wrong target by destroying the most instructive information in instructions at…
-
15 Jun 2021 1 repository listed Syntology ran 1 of 4 samples · 3 unverifiedThen, we cross-connect the key views of different scenes to construct augmented scenes.
-
19 Apr 2021 1 repository listedOne key challenge in this task is to ground instructions with the current visual information that the agent perceives.
-
9 Apr 2021 1 repository listedVision-and-Language Navigation (VLN) requires an agent to find a path to a remote location on the basis of natural-language instructions and a set of photo-realistic panoramas.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections