{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bayesian-batch-active-learning-as-sparse","title":"Bayesian Batch Active Learning as Sparse Subset Approximation","arxiv_id":"1908.02144","date":"2019-08-06","proceeding":"NeurIPS 2019 12","authors":["Robert Pinsler","Jonathan Gordon","Eric Nalisnick","José Miguel Hernández-Lobato"],"abstract":"Leveraging the wealth of unlabeled data produced in recent years provides great potential for improving supervised models. When the cost of acquiring labels is high, probabilistic active learning methods can be used to greedily select the most informative data points to be labeled. However, for many large-scale problems standard greedy procedures become computationally infeasible and suffer from negligible model change. In this paper, we introduce a novel Bayesian batch active learning approach that mitigates these issues. Our approach is motivated by approximating the complete data posterior of the model parameters. While naive batch construction methods result in correlated queries, our algorithm produces diverse batches that enable efficient active learning at scale. We derive interpretable closed-form solutions akin to existing active learning procedures for linear models, and generalize to arbitrary models using random projections. We demonstrate the benefits of our approach on several large-scale regression and classification tasks.","url_abs":"https://arxiv.org/abs/1908.02144v4","url_pdf":"https://arxiv.org/pdf/1908.02144v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bayesian-batch-active-learning-as-sparse","repo_url":"https://github.com/rpinsler/active-bayesian-coresets","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"bayesian-batch-active-learning-as-sparse","repo_url":"https://github.com/blackhc/active-bayesian-coresets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1908.02144","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1908.02144"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/blackhc/active-bayesian-coresets","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rpinsler/active-bayesian-coresets","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran_honours":2},"by_repo_kind":{"listed":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"b764c1f1d307c3fc","entry":"posterior","repo":"blackhc/active-bayesian-coresets","repo_kind":"listed","path":"experiments/linear_regression_active.py","file_url":"https://github.com/blackhc/active-bayesian-coresets/blob/HEAD/experiments/linear_regression_active.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"b764c1f1d307c3fc"}},{"code_sha256_prefix":"1985e7df2f54f090","entry":"posterior_tb","repo":"blackhc/active-bayesian-coresets","repo_kind":"listed","path":"experiments/linear_regression_active.py","file_url":"https://github.com/blackhc/active-bayesian-coresets/blob/HEAD/experiments/linear_regression_active.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"1985e7df2f54f090"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}