Papers › Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks

Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks

18 Nov 2019CVPR 2020 6arXiv:1911.07883archive 2025-07-28

Fengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan Liang

Vision-Language Navigation (VLN) is a task where agents learn to navigate following natural language instructions. The key to this task is to perceive both the visual scene and natural language sequentially. Conventional approaches exploit the vision and language features in cross-modal grounding. However, the VLN task remains challenging, since previous works have neglected the rich semantic information contained in the environment (such as implicit navigation graphs or sub-trajectory semantics). In this paper, we introduce Auxiliary Reasoning Navigation (AuxRN), a framework with four self-supervised auxiliary reasoning tasks to take advantage of the additional training signals derived from the semantic information. The auxiliary tasks have four reasoning objectives: explaining the previous actions, estimating the navigation progress, predicting the next orientation, and evaluating the trajectory consistency. As a result, these additional training signals help the agent to acquire knowledge of semantic representations in order to reason about its activity and build a thorough perception of the environment. Our experiments indicate that auxiliary reasoning tasks improve both the performance of the main task and the model generalizability by a large margin. Empirically, we demonstrate that an agent trained with self-supervised auxiliary reasoning tasks substantially outperforms the previous state-of-the-art method, being the best existing approach on the standard benchmark.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

NavigateVision-Language Navigation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Beam Search) error 3.24 #14 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Beam Search) length 40.85 #14 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Beam Search) oracle success 0.81 #14 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Beam Search) spl 0.21 #14 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Beam Search) success 0.71 #14 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Pre-explore) error 3.69 #28 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Pre-explore) length 10.43 #28 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Pre-explore) oracle success 0.75 #28 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Pre-explore) spl 0.65 #28 of 145 Archive leaderboard report
Vision and Language Navigation VLN Challenge Self-Supervised Auxiliary Reasoning Tasks (Pre-explore) success 0.68 #28 of 145 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections