Datasets › R2R

R2R (Room-to-Room)

Introduced by Peter Anderson et al. in Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments1 Jan 2018 archive 2025-07-28

R2R is a dataset for visually-grounded natural language navigation in real buildings. The dataset requires autonomous agents to follow human-generated navigation instructions in previously unseen buildings, as illustrated in the demo above. For training, each instruction is associated with a Matterport3D Simulator trajectory. 22k instructions are available, with an average length of 29 words. There is a test evaluation server for this dataset available at EvalAI.

Source: Natural language interaction with robots

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

13 shown of 13 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 174. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Agent Journey Beyond RGB: Unveiling Hybrid Semantic-Spatial Environmental Representations for Vision-and-Language Navigation 1 1 9 Dec 2024 not harvested
Towards Learning a Generalist Model for Embodied Navigation 2 1 4 Dec 2023 ran 10 of 15 samples (5 unverified)
VLN-PETL: Parameter-Efficient Transfer Learning for Vision-and-Language Navigation 1 1 20 Aug 2023 not harvested
Meta-Explore: Exploratory Hierarchical Vision-and-Language Navigation Using Scene Object Spectrum Grounding 0 1 7 Mar 2023 not harvested
BEVBert: Multimodal Map Pre-training for Language-guided Navigation 1 1 8 Dec 2022 not harvested
HOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation 1 1 22 Mar 2022 ran 4 of 12 samples (8 unverified)
Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation 1 1 23 Feb 2022 ran 4 of 5 samples (1 unverified; 5 pointer-only for licence)
A Recurrent Vision-and-Language BERT for Navigation 1 1 26 Nov 2020 not harvested
Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training 1 1 25 Feb 2020 ran 0 of 1 samples (1 unverified)
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout 1 1 8 Apr 2019 ran 0 of 4 samples (4 unverified)
Tactical Rewind: Self-Correction via Backtracking in Vision-and-Language Navigation 1 1 6 Mar 2019 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation 0 2 25 Nov 2018 not harvested
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments 8 1 20 Nov 2017 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only)

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Room2Room
  • R2R

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections