Datasets › Natural Questions › Papers where code ran, page 1
Natural Questions
Papers archive 2025-07-28
papers with a benchmark row: 49 · with a code link: 41 · where Syntology ran a sample: 25 (22 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument) Syntology
Show:
all papers with a benchmark row only where code ran (25 of 49 with a benchmark row: 22 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.
Page 1 of 1: papers 1 to 25 of the 25 papers with a benchmark row here where Syntology ran at least one harvested sample (22 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
A run counts when it is a run with no instrument failure any run, instrument failures included The first hides, in your browser, the papers where every run was a failure of Syntology's instrument; the second shows every paper on this page.
The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 1,404. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since 2026-09-28 , the first build that kept a record date for it; when this build read Syntology's graph is in the build record .
Paper Code Results Date Samples run Syntology
BM25S: Orders of magnitude faster lexical search via eager sparse scoring
3
4
4 Jul 2024
official (archive's flag): 10 ran · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
1
1
26 Jun 2024
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (3 pointer-only for licence )
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers
1
2
22 Mar 2024
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence )
Mistral 7B
6
1
10 Oct 2023
official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence )
Llama 2: Open Foundation and Fine-Tuned Chat Models
19
1
18 Jul 2023
community repositories only · 33 ran (of which 7 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 1 violated, 20 with no contract checked; 11 where Syntology's instrument failed) · 19 unverified (20 pointer-only for licence )
LLaMA: Open and Efficient Foundation Language Models
57
4
27 Feb 2023
official: no sample here; runs from other or unrecorded repositories · 37 ran (of which 9 constructed an object rather than computing a result; 25 with no instrument failure: 3 honoured, 0 violated, 22 with no contract checked; 12 where Syntology's instrument failed) · 21 unverified (4 pointer-only for licence )
REPLUG: Retrieval-Augmented Black-Box Language Models
3
2
30 Jan 2023
10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (13 pointer-only for licence )
Ask Me Anything: A simple strategy for prompting language models
3
3
5 Oct 2022
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
Atlas: Few-shot Learning with Retrieval Augmented Language Models
2
4
5 Aug 2022
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence )
PaLM: Scaling Language Modeling with Pathways
7
3
5 Apr 2022
32 ran (of which 16 constructed an object rather than computing a result; 24 with no instrument failure: 2 honoured, 1 violated, 21 with no contract checked; 8 where Syntology's instrument failed) · 5 unverified
Training Compute-Optimal Large Language Models
2
1
29 Mar 2022
8 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (4 pointer-only for licence )
SGPT: GPT Sentence Embeddings for Semantic Search
1
2
17 Feb 2022
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Improving language models by retrieving from trillions of tokens
2
1
8 Dec 2021
16 ran (of which 5 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 3 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (3 pointer-only for licence )
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering
2
2
9 Jun 2021
official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence )
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
3
2
17 Apr 2021
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
5
1
1 Jul 2020
official (archive's flag): 10 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence )
Generating Diverse and Consistent QA pairs from Contexts with Information-Maximizing Hierarchical Conditional VAEs
1
2
28 May 2020
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (3 pointer-only for licence )
Language Models are Few-Shot Learners
67
1
28 May 2020
community repositories only · 45 ran (of which 0 constructed an object rather than computing a result; 40 with no instrument failure: 2 honoured, 1 violated, 37 with no contract checked; 5 where Syntology's instrument failed) · 20 unverified (7 pointer-only for licence )
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
18
1
22 May 2020
6 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Dense Passage Retrieval for Open-Domain Question Answering
19
2
10 Apr 2020
official: harvested, nothing ran · 11 ran (of which 6 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (9 pointer-only for licence )
REALM: Retrieval-Augmented Language Model Pre-Training
6
1
10 Feb 2020
community repositories only · 4 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 2 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Reformer: The Efficient Transformer
10
1
13 Jan 2020
community repositories only · 6 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified
Generating Long Sequences with Sparse Transformers
7
1
23 Apr 2019
community repositories only · 5 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
A BERT Baseline for the Natural Questions
3
1
24 Jan 2019
community repositories only · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
Reading Wikipedia to Answer Open-Domain Questions
10
1
31 Mar 2017
community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence )