Browse State-of-the-Art › Visual Navigation
Visual Navigation
132 papers with code · 6 benchmarks · 19 datasets archive 2025-07-28
Visual Navigation is the problem of navigating an agent, e.g. a mobile robot, in an environment using camera input only. The agent is given a target image (an image it will see from the target position), and its goal is to move from its current position to the target by applying a sequence of actions, based on the camera observations only.
Source: Vision-based Navigation Using Deep Reinforcement Learning
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Cooperative Vision-and-Dialogue Navigation (19 rows) | NaviLLM | Towards Learning a Generalist Model for Embodied Navigation | code | Syntology ran 10 of 15 samples · 5 unverified | Compare |
| R2R (11 rows) | SUSA | Agent Journey Beyond RGB: Unveiling Hybrid Semantic-Spatial... | code | — | Compare |
| SOON Test (6 rows) | AutoVLN | Learning from Unlabeled 3D Environments for Vision-and-Language Navigation | code | — | Compare |
| AI2-THOR (2 rows) | MVV-IN | Multimodal Aggregation Approach for Memory Vision-Voice Indoor... | — | — | Compare |
| Dmlab-30 (1 row) | PopArt-IMPALA | Multi-task Deep Reinforcement Learning with PopArt | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| Help, Anna! (HANNA) (1 row) | Prevalent | Towards Learning a Generic Agent for Vision-and-Language... | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
19 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 132 papers with code (316 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Nov 2017 8 repositories listedThis is significant because a robot interpreting a natural-language navigation instruction on the basis of what it sees is carrying out a vision and language process that is similar to Visual Question Answering.
-
13 Feb 2017 6 repositories listedThe accumulated belief of the world enables the agent to track visited regions of the environment.
-
6 Jan 2020 4 repositories listed Syntology ran 1 of 6 samples · 5 unverifiedTo this end, we propose a new federated learning algorithm that jointly learns compact local representations on each device and a global model across all devices.
-
13 Dec 2019 4 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedSecond, we investigate the sim2real predictivity of Habitat-Sim for PointGoal navigation.
-
3 Mar 2023 3 repositories listedReal world applications of Reinforcement Learning (RL) are often partially observable, thus requiring memory.
-
5 Mar 2019 3 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedAs deep learning continues to make progress for challenging perception tasks, there is increased interest in combining vision, language, and decision-making.
-
15 May 2018 3 repositories listedWe propose to using high level semantic and contextual features including segmentation and detection masks obtained by off-the-shelf state-of-the-art vision as observations and use deep network to learn the navigation…
-
4 May 2018 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedAs part of our general methodology we discuss the software mapping techniques that enable the state-of-the-art deep convolutional neural network presented in [1] to be fully executed on-board within a strict 6 fps…
-
4 Dec 2023 2 repositories listed Syntology ran 10 of 15 samples · 5 unverifiedWe conduct extensive experiments to evaluate the performance and generalizability of our model.
-
26 May 2023 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 2 pointer-only (licence)Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling.
-
16 Jun 2022 2 repositories listedWe introduce SoundSpaces 2.
-
8 Aug 2021 2 repositories listedTo avoid the potentially hazardous trial-and-error of reinforcement learning, we focus on differentiable planners such as Value Iteration Networks (VIN), which are trained offline from safe expert demonstrations.
-
13 Jul 2021 2 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedIn the context of visual navigation, the capacity to map a novel environment is necessary for an agent to exploit its observation history in the considered place and efficiently reach known goals.
-
15 Feb 2021 2 repositories listedSpatial memory, or the ability to remember and recall specific locations and objects, is central to autonomous agents' ability to carry out tasks in real environments.
-
21 Oct 2020 2 repositories listedThis precludes the use of the learned policy on a real robot.
-
6 Aug 2020 2 repositories listedLearning a new task often requires both exploring to gather task-relevant information and exploiting this information to solve the task.
-
31 Dec 2019 2 repositories listedWhen training a neural network for a desired task, one may prefer to adapt a pre-trained network rather than starting from randomly initialized weights.
-
24 Dec 2019 2 repositories listedMoving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment.
-
18 Aug 2019 2 repositories listedIn this paper, we show how novel transfer reinforcement learning techniques can be applied to the complex task of target driven navigation using the photorealistic AI2THOR simulator.
-
10 Jul 2019 2 repositories listedTo train agents that search an environment for a goal location, we define the Navigation from Dialog History task.
-
10 May 2019 2 repositories listedNano-size unmanned aerial vehicles (UAVs), with few centimeters of diameter and sub-10 Watts of total power budget, have so far been considered incapable of running sophisticated visual-based autonomous navigation…
-
3 May 2019 2 repositories listedSelf-supervised learning aims to learn representations from the data itself without explicit manual supervision.
-
5 Mar 2019 2 repositories listedNumerous past works have tackled the problem of task-driven navigation.
-
10 Jan 2019 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThe Vision-and-Language Navigation (VLN) task entails an agent following navigational instruction in photo-realistic unknown environments.
-
3 Dec 2018 2 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedIn this paper we study the problem of learning to learn at both training and test time in the context of visual navigation.
-
16 Sep 2016 2 repositories listedTo address the second issue, we propose AI2-THOR framework, which provides an environment with high-quality 3D scenes and physics engine.
-
16 May 2025 1 repository listed Syntology ran 4 of 5 samples · 1 unverifiedRecent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse tasks.
-
25 Apr 2025 1 repository listedTo support the Low Altitude Economy (LAE), it is essential to achieve precise localization of unmanned aerial vehicles (UAVs) in urban areas where global positioning system (GPS) signals are unavailable.
-
14 Apr 2025 1 repository listedIn visual navigation, previous diffusion-based policies typically generate action sequences by initiating from denoising Gaussian noise.
-
14 Apr 2025 1 repository listedTo address these limitations, we propose a hybrid approach that combines the strengths of learning-based methods and classical approaches for RGB-only visual navigation.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections