{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluating-predictive-distributions-does-1","title":"The Neural Testbed: Evaluating Joint Predictions","arxiv_id":"2110.04629","date":"2021-10-09","proceeding":null,"authors":["Ian Osband","Zheng Wen","Seyed Mohammad Asghari","Vikranth Dwaracherla","Botao Hao","Morteza Ibrahimi","Dieterich Lawson","Xiuyuan Lu","Brendan O'Donoghue","Benjamin Van Roy"],"abstract":"Predictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open-source benchmark for controlled and principled evaluation of agents that generate such predictions. Crucially, the testbed assesses agents not only on the quality of their marginal predictions per input, but also on their joint predictions across many inputs. We evaluate a range of agents using a simple neural network data generating process. Our results indicate that some popular Bayesian deep learning agents do not fare well with joint predictions, even when they can produce accurate marginal predictions. We also show that the quality of joint predictions drives performance in downstream decision tasks. We find these results are robust across choice a wide range of generative models, and highlight the practical importance of joint predictions to the community.","url_abs":"https://arxiv.org/abs/2110.04629v4","url_pdf":"https://arxiv.org/pdf/2110.04629v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluating-predictive-distributions-does-1","repo_url":"https://github.com/deepmind/neural_testbed","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"evaluating-predictive-distributions-does-1","repo_url":"https://github.com/google-deepmind/neural_testbed","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2110.04629","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2110.04629"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google-deepmind/neural_testbed","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/deepmind/neural_testbed","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":2},"by_repo_kind":{"listed":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"467b5f8cab99e0d6","entry":"combine_leaderboards","repo":"google-deepmind/neural_testbed","repo_kind":"listed","path":"neural_testbed/leaderboard/score.py","file_url":"https://github.com/google-deepmind/neural_testbed/blob/HEAD/neural_testbed/leaderboard/score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"467b5f8cab99e0d6"}},{"code_sha256_prefix":"d3c792b894c4813b","entry":"logging_freq","repo":"google-deepmind/neural_testbed","repo_kind":"listed","path":"neural_testbed/agents/enn_agent.py","file_url":"https://github.com/google-deepmind/neural_testbed/blob/HEAD/neural_testbed/agents/enn_agent.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d3c792b894c4813b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}