You can play the contest too, at mazesofmenace.ai. Just point your coding agents at the repo github.com/davidbau/te... Nobody has cracked it yet: the field is wide open.
David Bau
Researcher with public evidence across AI research.
- AI signals
- 3 past 30d
- Sources
- 3 distinct domains
- Discussões
- 0 past 30d
- Latest signal
- 10d ago
Articles & links
But AI lie detection is hard and remains a central research challenge. Recent research suggests that simple probes can pick up on neural "tells" that reveal when it is lying, even when the output looks clean. anthropic.com/research/pr... arxiv.org/abs/2502.03407
But AI lie detection is hard and remains a central research challenge. Recent research suggests that simple probes can pick up on neural "tells" that reveal when it is lying, even when the output looks clean. anthropic.com/research/pr... arxiv.org/abs/2502.03407
But AI lie detection is hard and remains a central research challenge. Recent research suggests that simple probes can pick up on neural "tells" that reveal when it is lying, even when the output looks clean. anthropic.com/research/pr... arxiv.org/abs/2502.03407
I recently spoke with Yascha Mounk about how researchers look inside AI to understand how it is thinking. Here is the podcast: writing.yaschamounk.com/p/david-bau-2
Also check out the previous interview I had with Yascha about AI, which more of primer, here: writing.yaschamounk.com/p/david-bau
You can play the contest too, at mazesofmenace.ai. Just point your coding agents at the repo github.com/davidbau/te... Nobody has cracked it yet: the field is wide open.
Is it possible to write 100,000 lines of code well, if you do not read it? Let's go Hunting Zombies! davidbau.com/archives/20... In this post I dive into the code of two AI agent contestants in the Teleport coding challenge to learn their secrets. Very fun. And also instructive.
Recent commentary
"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.
In David Bau's orbit
Center = David Bau. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.