Datasets › Natural Questions › Papers where code ran, page 1

Natural Questions

Papers archive 2025-07-28

papers with a benchmark row: 49 · with a code link: 41 · where Syntology ran a sample: 25 (22 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (25 of 49 with a benchmark row: 22 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument)

Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.

Page 1 of 1: papers 1 to 25 of the 25 papers with a benchmark row here where Syntology ran at least one harvested sample (22 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 1,404. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
BM25S: Orders of magnitude faster lexical search via eager sparse scoring 3 4 4 Jul 2024 official (archive's flag): 10 ran · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation 1 1 26 Jun 2024 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (3 pointer-only for licence)
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers 1 2 22 Mar 2024 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
Mistral 7B 6 1 10 Oct 2023 official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 1 18 Jul 2023 community repositories only · 33 ran (of which 7 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 1 violated, 20 with no contract checked; 11 where Syntology's instrument failed) · 19 unverified (20 pointer-only for licence)
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 official: no sample here; runs from other or unrecorded repositories · 37 ran (of which 9 constructed an object rather than computing a result; 25 with no instrument failure: 3 honoured, 0 violated, 22 with no contract checked; 12 where Syntology's instrument failed) · 21 unverified (4 pointer-only for licence)
REPLUG: Retrieval-Augmented Black-Box Language Models 3 2 30 Jan 2023 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (13 pointer-only for licence)
Ask Me Anything: A simple strategy for prompting language models 3 3 5 Oct 2022 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
Atlas: Few-shot Learning with Retrieval Augmented Language Models 2 4 5 Aug 2022 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 32 ran (of which 16 constructed an object rather than computing a result; 24 with no instrument failure: 2 honoured, 1 violated, 21 with no contract checked; 8 where Syntology's instrument failed) · 5 unverified
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 8 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (4 pointer-only for licence)
SGPT: GPT Sentence Embeddings for Semantic Search 1 2 17 Feb 2022 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Improving language models by retrieving from trillions of tokens 2 1 8 Dec 2021 16 ran (of which 5 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 3 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (3 pointer-only for licence)
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering 2 2 9 Jun 2021 official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence)
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models 3 2 17 Apr 2021 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval 5 1 1 Jul 2020 official (archive's flag): 10 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence)
Generating Diverse and Consistent QA pairs from Contexts with Information-Maximizing Hierarchical Conditional VAEs 1 2 28 May 2020 official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (3 pointer-only for licence)
Language Models are Few-Shot Learners 67 1 28 May 2020 community repositories only · 45 ran (of which 0 constructed an object rather than computing a result; 40 with no instrument failure: 2 honoured, 1 violated, 37 with no contract checked; 5 where Syntology's instrument failed) · 20 unverified (7 pointer-only for licence)
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks 18 1 22 May 2020 6 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Dense Passage Retrieval for Open-Domain Question Answering 19 2 10 Apr 2020 official: harvested, nothing ran · 11 ran (of which 6 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (9 pointer-only for licence)
REALM: Retrieval-Augmented Language Model Pre-Training 6 1 10 Feb 2020 community repositories only · 4 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 2 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Reformer: The Efficient Transformer 10 1 13 Jan 2020 community repositories only · 6 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified
Generating Long Sequences with Sparse Transformers 7 1 23 Apr 2019 community repositories only · 5 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
A BERT Baseline for the Natural Questions 3 1 24 Jan 2019 community repositories only · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
Reading Wikipedia to Answer Open-Domain Questions 10 1 31 Mar 2017 community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)