Simple probes can catch sleeper agents
3 experts across 2 network communities independently surfaced this.
Research & technical analysis
2 expertsEvidence, methods and technical implications.
“But AI lie detection is hard and remains a central research challenge. Recent research suggests that simple probes can pick up on neural "tells" that reveal when it is lying, even when the output looks clean. anthropic.com/research/pr... arxiv.org/abs/2502.…”