{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2603-06271","title":"Agentic retrieval-augmented reasoning reshapes collective reliability under model variability in radiology question answering","arxiv_id":"2603.06271","date":"2026-03-06","proceeding":null,"authors":["Mina Farajiamiri","Jeta Sopa","Saba Afza","Lisa Adams","Felix Barajas Ordonez","Tri-Thien Nguyen","Mahshad Lotfinia","Sebastian Wind","Keno Bressem","Sven Nebelung","Daniel Truhn","Soroosh Tayebi Arasteh"],"abstract":"Agentic retrieval-augmented reasoning pipelines are increasingly used to structure how large language models (LLMs) incorporate external evidence in clinical decision support. These systems iteratively retrieve curated domain knowledge and synthesize it into structured reports before answer selection. Although such pipelines can improve performance, their impact on reliability under model variability remains unclear. In real-world deployment, heterogeneous models may align, diverge, or synchronize errors in ways not captured by accuracy. We evaluated 34 LLMs on 169 expert-curated publicly available radiology questions, comparing zero-shot inference with a radiology-specific multi-step agentic retrieval condition in which all models received identical structured evidence reports derived from curated radiology knowledge. Agentic inference reduced inter-model decision dispersion (median entropy 0.48 vs. 0.13) and increased robustness of correctness across models (mean 0.74 vs. 0.81). Majority consensus also increased overall (P<0.001). Consensus strength and robust correctness remained correlated under both strategies (\\r{ho}=0.88 for zero-shot; \\r{ho}=0.87 for agentic), although high agreement did not guarantee correctness. Response verbosity showed no meaningful association with correctness. Among 572 incorrect outputs, 72% were associated with moderate or high clinically assessed severity, although inter-rater agreement was low (\\k{appa}=0.02). Agentic retrieval therefore was associated with more concentrated decision distributions, stronger consensus, and higher cross-model robustness of correctness. These findings suggest that evaluating agentic systems through accuracy or agreement alone may not always be sufficient, and that complementary analyses of stability, cross-model robustness, and potential clinical impact are needed to characterize reliability under model variability.","url_abs":"https://arxiv.org/abs/2603.06271","url_pdf":"https://arxiv.org/pdf/2603.06271","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2603.06271","atlas_url":"https://app.syntology.ai/?focus=2603.06271","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2603.06271"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/vllm-project/vllm","reach":null},{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/sopajeta/RaR","reach":null}],"summary":{"unverified":6},"by_repo_kind":{"found_in_text":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"e9c17abfaf588d78","entry":"deduplicate_and_format_sources","repo":"sopajeta/RaR","repo_kind":"found_in_text","path":"agentic_retrieval/utils.py","file_url":"https://github.com/sopajeta/RaR/blob/HEAD/agentic_retrieval/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e9c17abfaf588d78"}},{"code_sha256_prefix":"ca8747f51ea457e7","entry":"get_config_value","repo":"sopajeta/RaR","repo_kind":"found_in_text","path":"agentic_retrieval/utils.py","file_url":"https://github.com/sopajeta/RaR/blob/HEAD/agentic_retrieval/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"ca8747f51ea457e7"}},{"code_sha256_prefix":"456e5966a0a86bdd","entry":"get_search_params","repo":"sopajeta/RaR","repo_kind":"found_in_text","path":"agentic_retrieval/utils.py","file_url":"https://github.com/sopajeta/RaR/blob/HEAD/agentic_retrieval/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"456e5966a0a86bdd"}},{"code_sha256_prefix":"5aa31dfe33a58449","entry":"is_bad_report","repo":"sopajeta/RaR","repo_kind":"found_in_text","path":"agentic_retrieval/dataset.py","file_url":"https://github.com/sopajeta/RaR/blob/HEAD/agentic_retrieval/dataset.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"5aa31dfe33a58449"}},{"code_sha256_prefix":"e95d03854ee33187","entry":"load_existing_from_json","repo":"sopajeta/RaR","repo_kind":"found_in_text","path":"agentic_retrieval/persistence.py","file_url":"https://github.com/sopajeta/RaR/blob/HEAD/agentic_retrieval/persistence.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e95d03854ee33187"}},{"code_sha256_prefix":"f819cb72cb280c2a","entry":"load_raw_questions","repo":"sopajeta/RaR","repo_kind":"found_in_text","path":"agentic_retrieval/dataset.py","file_url":"https://github.com/sopajeta/RaR/blob/HEAD/agentic_retrieval/dataset.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"f819cb72cb280c2a"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 7,081 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9623,"papers_checked":7081},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}