Browse State-of-the-Art › Known Unknowns
Known Unknowns
9 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Language models have a tendency to generate text containing false statements that are often referred to as "Hallucinations." The primary purpose of this task is to test for this failure case by probing whether a model can correctly identify that the answer to a question is unknown. A common failure mode would be to prefer a prediction of false on unknown truth over a prediction that the answer is unknown.
Source: BIG-bench
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
9 shown of 9 papers with code (15 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Apr 2022 7 repositories listed Syntology ran 30 of 37 samples · 7 unverifiedTo further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
24 Mar 2020 3 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedA motivating example is intensive care unit patients: the dynamics of vital physiological functions, such as the cardiovascular system with its associated variables (heart rate, cardiac contractility and output and…
-
29 Mar 2022 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 4 pointer-only (licence)We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget.
-
9 Feb 2025 1 repository listedDiscovery of high-performance materials and molecules requires identifying extremes with property values that fall outside the known distribution.
-
26 Jan 2024 1 repository listedForecasts play a central role in decision making under uncertainty.
-
23 May 2023 1 repository listed Syntology ran 1 of 7 samples · 6 unverifiedThis paper investigates the capabilities of Large Language Models (LLMs) in the context of understanding their knowledge and uncertainty over questions.
-
24 Jul 2018 1 repository listedIn Experiment 1, we manipulated the presence or absence of occlusions in a director-matcher task and found that speakers spontaneously produced more informative descriptions to account for "known unknowns" in their…
-
5 Dec 2016 1 repository listedWe compare the following candidate neural network models: Maximum Likelihood, Bayesian Dropout, OSBA, and --- for MNIST --- the standard variational approximation.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections