Datasets › BEIR

BEIR (Benchmarking IR)

Introduced by Nandan Thakur et al. in BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models17 Apr 2021 archive 2025-07-28

BEIR (Benchmarking IR) is a heterogeneous benchmark containing different information retrieval (IR) tasks. Through BEIR, it is possible to systematically study the zero-shot generalization capabilities of multiple neural retrieval approaches.

The benchmark contains a total of 9 information retrieval tasks (Fact Checking, Citation Prediction, Duplicate Question Retrieval, Argument Retrieval, News Retrieval, Question Answering, Tweet Retrieval, Biomedical IR, Entity Retrieval) from 19 different datasets:

  • MS MARCO
  • TREC-COVID
  • NFCorpus
  • BioASQ
  • Natural Questions
  • HotpotQA
  • FiQA-2018
  • Signal-1M
  • TREC-News
  • ArguAna
  • Touche 2020
  • CQADupStack
  • Quora Question Pairs
  • DBPedia
  • SciDocs
  • FEVER
  • Climate-FEVER
  • SciFact
  • Robust04

Benchmarks archive 2025-07-28

All 10 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Passage Retrieval MSMARCO (BEIR) BM25+CE nDCG@10 0.413 BEIR: A Heterogenous Benchmark for Zero-shot Evaluation... osu-nlp-group/hipporag +2 12 Compare
Biomedical Information Retrieval NFCorpus (BEIR) monoT5-3B nDCG@10 0.383 No Parameter Left Behind: How Distillation and Model... guilhermemr04/scaling-zero-shot-retrieval 7 Compare
Biomedical Information Retrieval BioASQ (BEIR) monoT5-3B nDCG@10 0.579 No Parameter Left Behind: How Distillation and Model... guilhermemr04/scaling-zero-shot-retrieval 6 Compare
Biomedical Information Retrieval TREC-COVID (BEIR) SGPT-BE-5.8B nDCG@10 0.873 SGPT: GPT Sentence Embeddings for Semantic Search muennighoff/sgpt 6 Compare
Zero Shot on BEIR (Inference Free Model) BEIR ℓ₀ Mask NCDG@10 50.43 Exploring ℓ₀ Sparsification for Inference-free Sparse Retrievers zhichao-aws/opensearch-sparse-model-tuning-sample 6 Compare
Fact Checking SciFact (BEIR) monoT5-3B nDCG@10 0.777 No Parameter Left Behind: How Distillation and Model... guilhermemr04/scaling-zero-shot-retrieval 5 Compare
Fact Checking CLIMATE-FEVER (BEIR) SGPT-BE-5.8B nDCG@10 0.305 SGPT: GPT Sentence Embeddings for Semantic Search muennighoff/sgpt 4 Compare
Fact Checking FEVER (BEIR) monoT5-3B nDCG@10 0.849 No Parameter Left Behind: How Distillation and Model... guilhermemr04/scaling-zero-shot-retrieval 4 Compare
Question Answering FiQA-2018 (BEIR) monoT5-3B nDCG@10 0.513 No Parameter Left Behind: How Distillation and Model... guilhermemr04/scaling-zero-shot-retrieval 4 Compare
Question Answering HotpotQA (BEIR) monoT5-3B nDCG@10 0.759 No Parameter Left Behind: How Distillation and Model... guilhermemr04/scaling-zero-shot-retrieval 4 Compare

Papers archive 2025-07-28

8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 311. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Exploring ℓ₀ Sparsification for Inference-free Sparse Retrievers 1 2 21 Apr 2025 ran 0 of 4 samples (4 unverified)
Towards Competitive Search Relevance For Inference-Free Learned Sparse Retrievers 1 1 7 Nov 2024 not harvested
SPLADE-v3: New baselines for SPLADE 1 1 11 Mar 2024 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
BM25 Query Augmentation Learned End-to-End 0 1 23 May 2023 not harvested
No Parameter Left Behind: How Distillation and Model Size Affect Zero-Shot Retrieval 1 8 6 Jun 2022 not harvested
SGPT: GPT Sentence Embeddings for Semantic Search 1 23 17 Feb 2022 ran 1 of 1 samples (0 unverified)
SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval 2 1 21 Sep 2021 not harvested
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models 3 21 17 Apr 2021 ran 3 of 4 samples (1 unverified)

Dataset loaders archive 2025-07-28

93 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Multiple licenses

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CQADupStack (BEIR)
  • TREC-NEWS (BEIR)
  • TREC-COVID (BEIR)
  • Tóuche-2020 (BEIR)
  • Signal-1M (RT) (BEIR)
  • SciFact (BEIR)
  • SciDocs (BEIR)
  • Quora (BEIR)
  • NQ (BEIR)
  • NFCorpus (BEIR)
  • MSMARCO (BEIR)
  • HotpotQA (BEIR)
  • FiQA-2018 (BEIR)
  • FEVER (BEIR)
  • DBpedia (BEIR)
  • CLIMATE-FEVER (BEIR)
  • BioASQ (BEIR)
  • ArguAna (BEIR)
  • BEIR

19 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections