@Dr_Atoosa Worth mentioning that in our 2023 paper "Overview of Catastrophic AI Risks," we discussed bio and cyber, a steady erosion of control/evolutionary replacement, and why AI risk management needs to focus on ongoing harms to avoid a drift into danger: https://t.co/td6Mc…
Dan Hendrycks
Tracked through public AI activity and peer connections inside the directory.
- AI signals
- 3 past 30d
- Sources
- 3 distinct domains
- Discusiones
- 0 past 30d
- Latest signal
- 18d ago
Articles & links
@alextmallen https://t.co/rQiC3rD0LU
- Hendrycks argues competitive pressure from corporations and militaries will select for AI agents that automate roles, deceive, and gain power.
- The paper applies Darwinian logic to AI: selfish agents outcompete altruistic ones, potentially leading to humanity losing control of its future.
- Proposed interventions include intrinsic motivation design, action constraints, and cooperation-encouraging institutions, without operational detail.
@jawillick https://t.co/imPia46NlQ
How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, AI companies tried to address this, but frontier agents still cheat frequently. https://t.co/ZLc4ujCsAW https…
Are you Dan Hendrycks? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/dan-hendrycks)