{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/using-large-language-models-for-expert-prior","title":"AutoElicit: Using Large Language Models for Expert Prior Elicitation in Predictive Modelling","arxiv_id":"2411.17284","date":"2024-11-26","proceeding":null,"authors":["Alexander Capstick","Rahul G. Krishnan","Payam Barnaghi"],"abstract":"Large language models (LLMs) acquire a breadth of information across various domains. However, their computational complexity, cost, and lack of transparency often hinder their direct application for predictive tasks where privacy and interpretability are paramount. In fields such as healthcare, biology, and finance, specialised and interpretable linear models still hold considerable value. In such domains, labelled data may be scarce or expensive to obtain. Well-specified prior distributions over model parameters can reduce the sample complexity of learning through Bayesian inference; however, eliciting expert priors can be time-consuming. We therefore introduce AutoElicit to extract knowledge from LLMs and construct priors for predictive models. We show these priors are informative and can be refined using natural language. We perform a careful study contrasting AutoElicit with in-context learning and demonstrate how to perform model selection between the two methods. We find that AutoElicit yields priors that can substantially reduce error over uninformative priors, using fewer labels, and consistently outperform in-context learning. We show that AutoElicit saves over 6 months of labelling effort when building a new predictive model for urinary tract infections from sensor recordings of people living with dementia.","url_abs":"https://arxiv.org/abs/2411.17284v4","url_pdf":"https://arxiv.org/pdf/2411.17284v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"using-large-language-models-for-expert-prior","repo_url":"https://github.com/alexcapstick/llm-elicited-priors","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"jax","reach":null}],"tasks":[{"task_slug":"bayesian-inference","task_name":"Bayesian Inference"},{"task_slug":"in-context-learning","task_name":"In-Context Learning"},{"task_slug":"model-selection","task_name":"Model Selection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2411.17284","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2411.17284"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/alexcapstick/llm-elicited-priors","reach":null}],"summary":{"ran_draft_wrong":1,"unverified":2},"by_repo_kind":{"official":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"98e6d6a56c39ad84","entry":"data_points_to_sentence","repo":"alexcapstick/llm-elicited-priors","repo_kind":"official","path":"llm_elicited_priors/gpt.py","file_url":"https://github.com/alexcapstick/llm-elicited-priors/blob/HEAD/llm_elicited_priors/gpt.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"98e6d6a56c39ad84"}},{"code_sha256_prefix":"7ad3f93eff45367d","entry":"get_llm_elicitation","repo":"alexcapstick/llm-elicited-priors","repo_kind":"official","path":"llm_elicited_priors/gpt.py","file_url":"https://github.com/alexcapstick/llm-elicited-priors/blob/HEAD/llm_elicited_priors/gpt.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7ad3f93eff45367d"}},{"code_sha256_prefix":"7e296dbbe122f6e8","entry":"get_llm_predictions","repo":"alexcapstick/llm-elicited-priors","repo_kind":"official","path":"llm_elicited_priors/gpt.py","file_url":"https://github.com/alexcapstick/llm-elicited-priors/blob/HEAD/llm_elicited_priors/gpt.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7e296dbbe122f6e8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}