Papers › Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector...

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

24 Sep 2026arXiv:2609.29429added by Syntology

Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang

Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.

PaperPDF

In Syntology View this paper on Syntology, its page in Syntology's graph. That page lists the repositories linked to the paper, the abstract and the calls for agents.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, on Syntology's MCP service (how to connect):

Code

sumleo/RLCDAlignBench found in paper text by SyntologySyntology: not harvested report

Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Syntology holds the repository link but has not harvested or run code from it.

Results from the paper

The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2609.29429, was issued in September 2026, after that date, so the archive has no leaderboard rows for it.

Placed on leaderboards by Syntology Syntology

Syntology's extractor, the model Claude Sonnet 4.5, judged this paper's tables to report on 1 leaderboard (a model's judgement, not a result) and has not placed the paper on it: Summarization Consistency Evaluation · AggreFact (refused by a rule: the row label the model quoted is not the row's label). What is not shown.

Syntology has checked 8,886 of the 9,662 papers on this site that are newer than the archive (for 3,717 of them no archive leaderboard matched the paper's tables, so there was nothing further to check); 733 were read and have nothing to place (arXiv has no HTML version of the paper, or that version has no tables), 41 could not be read (the extractor's reply could not be parsed), and results from the other 2 appear after they are checked.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections