Home › Datasets › task › Vision-Language Navigation

Vision-Language Navigation datasets

archive 2025-07-28

9 datasets carry the task tag "Vision-Language Navigation" (the task itself: Vision-Language Navigation), ordered by the archive's paper count. Page 1 of 1: 9 shown of 9. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Vision-Language Navigation datasets 1–9 of 9

R2R (Room-to-Room)
R2R is a dataset for visually-grounded natural language navigation in real buildings.
174 papers · 2 benchmarks
TEACh (Task-driven Embodied Agents that Chat)
Robots operating in human spaces must be able to engage in natural language interaction with people, both understanding and executing instructions, and using conversation to resolve ambiguity and recover from mistakes.
36 papers · 0 benchmarks
Talk The Walk is a large-scale dialogue dataset grounded in action and perception.
11 papers · 0 benchmarks
ReaSCAN (ReaSCAN: Compositional Reasoning in Language Grounding)
ReaSCAN is a synthetic navigation task that requires models to reason about surroundings over syntactically difficult languages.
10 papers · 0 benchmarks
BnB is a large-scale and diverse in-domain VLN (Vision and Language Navigation) dataset.
3 papers · 0 benchmarks
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
XL-R2R (Cross-lingual Room-to-Room)
The XL-R2R dataset is built upon the R2R dataset and extends it with Chinese instructions.
2 papers · 0 benchmarks
PInNED (Personalized Instance-based Navigation Embodied Dataset)
In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly.
1 paper · 0 benchmarks
realfred is an embodied instruction following benchmark.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.