“3. Our recent ROGUE benchmark: our empirical test of how corrigible today’s frontier AI agents really are. Models that behave well in ordinary chat don’t always stay that way once they’re given a real off-switch they could disable to finish a task. Paper: a…”
“2. Corrigibility - building AI systems that stay editable, deferential, and willing to be shut down - is feasible with the lexicographic approach that sets hierarchical priorities on an agent’s objectives. Paper: arxiv.org/abs/2507.20964”
“In it, we discussed: 1. Why aligning AI to all human values is intractable — and what smaller, universal target we can aim for instead. Paper: arxiv.org/abs/2502.05934”
“There was no recording, but please check out @reecedkeller.bsky.social's talk on our work at our #Cosyne2026 "NeuroAgents" workshop: www.youtube.com/watch?v=LZiR... Slides here: anayebi.github.io/files/slides...”
“There was no recording, but please check out @reecedkeller.bsky.social's talk on our work at our #Cosyne2026 "NeuroAgents" workshop: www.youtube.com/watch?v=LZiR... Slides here: anayebi.github.io/files/slides...”
“In case you want to learn more about AI safety this 4th, check out the recent recording of some of my group's work on the AI Safety Research directory! www.youtube.com/watch?v=XWQJ...”
We use essential cookies to keep the site working (login, form security). With your permission, we also use analytics cookies to understand how you use the site.
Privacy policy