{"url":"/dataset/openeqa","name":"OpenEQA","full_name":null,"description_markdown":"The **OpenEQA** dataset is a significant contribution in the field of **Embodied Question Answering (EQA)**. Let me provide you with some details:\r\n\r\n1. **Definition**:\r\n   - **Embodied Question Answering (EQA)** involves understanding an environment well enough to answer questions about it in natural language.\r\n   - EQA agents can achieve this understanding through either **episodic memory** (as seen in agents using smart glasses) or **active exploration** of the environment (as in the case of mobile robots).\r\n\r\n2. **OpenEQA Dataset**:\r\n   - **OpenEQA** is the **first open-vocabulary benchmark dataset** for EQA that supports both episodic memory and active exploration use cases.\r\n   - It contains over **1600 high-quality human-generated questions** drawn from more than **180 real-world environments**.\r\n   - The dataset consists of **question-answer pairs** ($Q, A^*$) and corresponding **episode histories** ($H$).\r\n   - You can find the question-answer pairs in the file `data/open-eqa-v0.json`.\r\n   - To access the episode histories, follow the instructions provided [here](https://github.com/facebookresearch/open-eqa#dataset).\r\n\r\n3. **Evaluation Protocol**:\r\n   - OpenEQA also provides an **automatic evaluation protocol** powered by **language model-based evaluation** (LLM).\r\n   - This evaluation protocol correlates well with human judgment.\r\n\r\n4. **Foundation Models Evaluation**:\r\n   - Researchers evaluated several state-of-the-art foundation models, including **GPT-4V**, using the OpenEQA dataset.\r\n   - The findings revealed that these models significantly lag behind **human-level performance** in EQA tasks.\r\n\r\n5. **Significance**:\r\n   - OpenEQA serves as a **straightforward, measurable, and practically relevant benchmark** for current-generation foundation models.\r\n   - It poses a considerable challenge and inspires research at the intersection of **Embodied AI**, **conversational agents**, and **world models**.","description_withheld":null,"homepage":"https://open-eqa.github.io","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[{"name":"Embodied Question Answering","url":"/task/embodied-question-answering","datasets_with_task":"/datasets/task/embodied-question-answering"}],"languages":[],"variants":["OpenEQA"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}