Browse State-of-the-Art › TriviaQA
TriviaQA
60 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| BIG-bench (1 row) | Gopher-280B (few-shot, k=64) | Scaling Language Models: Methods, Analysis & Insights from Training Gopher | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 60 papers with code (124 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
10 Apr 2020 22 repositories listed Syntology ran 15 of 35 samples · 20 unverified · 5 pointer-only (licence)To address this limitation, we introduce the Longformer with an attention mechanism that scales linearly with sequence length, making it easy to process documents of thousands of tokens or longer.
-
2 Jul 2020 8 repositories listedGenerative models for open domain question answering have proven to be competitive, without resorting to external knowledge.
-
10 Nov 2019 7 repositories listed Syntology ran 4 of 6 samples · 2 unverified · 4 pointer-only (licence)We introduce an approach for open-domain question answering (QA) that retrieves and reads a passage graph, where vertices are passages of text and edges represent relationships that are derived from an external…
-
1 Jul 2020 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn much recent work, the retriever is a learned component that uses coarse-grained vector representations of questions and passages.
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
9 May 2017 3 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedWe present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples.
-
31 May 2024 2 repositories listedTo calibrate both implicit and explicit confidence markers, we introduce a pragmatic, listener-aware finetuning method (LACIE) that models the listener, considering not only whether an answer is right, but whether it…
-
2 Jan 2021 2 repositories listedWe also explore two approaches for end-to-end supervised training of the reader and retriever components in OpenQA models.
-
26 May 2025 1 repository listedSubsequently, we introduce a novel knowledge integration model that incorporates the retrieval knowledge into instructions during fine-tuning to intensify the model.
-
21 Dec 2024 1 repository listedThis paper proposes a novel approach to develop an open-domain and long-form Over-The-Top (OTT) Question-Answering (QA) dataset, DragonVerseQA, specifically oriented to the fantasy universe of "House of the Dragon" and…
-
10 Oct 2024 1 repository listedIn our method, a small auxiliary model is used to process the prompt and produce an approximation of the KV cache used by a base model.
-
24 Sep 2024 1 repository listedWe demonstrate that hints enhance the accuracy of answers more than retrieved and generated contexts.
-
12 Aug 2024 1 repository listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)Open Domain Question Answering (ODQA) has been advancing rapidly in recent times, driven by significant developments in dense passage retrieval and pretrained language models.
-
27 Jun 2024 1 repository listedOur study highlights the potential of finetuning on synthetic data for improving the performance of LLMs on longer-context tasks.
-
18 Jun 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Lastly, our research rediscovers the importance of using alignment metrics beyond simple percent alignment, showing that judges with high percent agreement can still assign vastly different scores.
-
17 Jun 2024 1 repository listedTo address this issue, we explore the task of "credibility-aware RAG", in which LLMs automatically adjust the influence of retrieved documents based on their credibility scores to counteract misinformation.
-
9 Jun 2024 1 repository listedThe Retrieval Augmented Generation (RAG) framework utilizes a combination of parametric knowledge and external knowledge to demonstrate state-of-the-art performance on open-domain question answering tasks.
-
4 Jun 2024 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)This decline can be attributed to the loss of key information during the compression process.
-
26 May 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedOpen-domain question answering (Open-QA) is a common task for evaluating large language models (LLMs).
-
25 Apr 2024 1 repository listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)We present LayerSkip, an end-to-end solution to speed-up inference of large language models (LLMs).
-
3 Apr 2024 1 repository listedIn Open-domain Question Answering (ODQA), it is essential to discern relevant contexts as evidence and avoid spurious ones among retrieved results.
-
8 Mar 2024 1 repository listedOpen-domain question answering (ODQA) has emerged as a pivotal research spotlight in information systems.
-
8 Mar 2024 1 repository listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)Leveraging our previous observations on controlling hallucinations, we propose an approach for learning more reliable reward models, and show that they improve the efficacy of RL factuality finetuning in long-form…
-
9 Oct 2023 1 repository listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)For example, natural language approaches cannot be transferred to image generation.
-
21 Jul 2023 1 repository listedOpen-domain question answering (QA) tasks usually require the retrieval of relevant information from a large corpus to generate accurate answers.
-
26 May 2023 1 repository listedOpen-Domain Question Answering (ODQA) systems necessitate a reader model capable of generating answers by simultaneously referring to multiple passages.
-
26 May 2023 1 repository listedThe Open-Domain Question Answering (ODQA) task involves retrieving and subsequently generating answers from fine-grained relevant passages within a database.
-
24 May 2023 1 repository listed Syntology ran 1 of 13 samples · 12 unverifiedWith the advance of large language models (LLMs), the research field of LLM applications becomes more and more popular and the idea of constructing pipelines to accomplish complex tasks by stacking LLM API calls come…
-
24 May 2023 1 repository listedA trustworthy real-world prediction system should produce well-calibrated confidence scores; that is, its confidence in an answer should be indicative of the likelihood that the answer is correct, enabling deferral to…
-
19 May 2023 1 repository listed Syntology ran 4 of 5 samples · 1 unverifiedUnlike these models, humans typically utilize external tools to cross-check and refine their initial content, like using a search engine for fact-checking, or a code interpreter for debugging.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections