huggingface.co web signal

'Just Ask Jev' Paper Reports 0.886 Median AUROC Detecting 10 LLM Alignment Failures at 63x Lower Cost

Safety Open Source ai-business

Summary

A new paper introduces Jev, a decision model trained with Reinforcement Learning for Calibrated Decisions to detect ten alignment-failure modes — sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty and power seeking — across 44 benchmarks and five target models. The authors report a median zero-shot AUROC of 0.886, matching human agreement with reference scorers, at roughly 63x lower cost than LLM-judge alternatives. Alongside the model, they release RLCDAlignBench as an evaluation suite.