Papers › Authority-Preserving Evaluation of Medical Vision-Language Assistants

Authority-Preserving Evaluation of Medical Vision-Language Assistants

14 Sep 2026arXiv:2609.22302added by Syntology

Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong, Roger Wattenhofer

Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).

Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local service retains authority to accept or replace that proposal under referral policy, capacity, and locally held patient context. Proposal quality and selected-action quality are therefore distinct evaluation targets, and benchmark evidence transfers between them only when local review preserves the expected action score. We introduce AuthEval, a logging and evaluation framework that records both actions, scores the selected action under declared local criteria, and, where feasible, scores the declined proposal under the same rule. It reports the resulting authority gap only when the record supports it. Because the gap is the product of the proposal-change rate and the mean score change on changed cases, that rate alone determines neither its magnitude nor its sign. On ISIC 2019, with MedGemma and simulated local review, two constraint regimes with similar change rates produced an optimistic image-equal gap under capacity (+0.744 simulator units) but no detectable gap under safety. The declared evaluation unit also mattered: under mixed constraints the gap reversed from +0.374 to -0.206 when weighting shifted from image to lesion-aware cluster. AuthEval thus clarifies whether a study's records support claims about the model, the workflow, or both.

PaperPDF

In Syntology View this paper on Syntology, its page in Syntology's graph. That page lists the repositories linked to the paper, the abstract and the calls for agents.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, on Syntology's MCP service (how to connect):

Code

flint-xf-fan/AuthEval found in paper text by SyntologySyntology: not harvested report

Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Syntology holds the repository link but has not harvested or run code from it.

Results from the paper

The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2609.22302, was issued in September 2026, after that date, so the archive has no leaderboard rows for it.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections