Jakub (OAI Chief Scientist) posted openai.com/index/an-ali.... I agree w/ many points. We should pause AI progress. Labs should voluntarily pause in absence of coordination. Thank you Jakub! World is counting on you (+ other OAI employees) to keep OAI honest to these words.
Alex Turner
Researcher with public evidence across AI research, Models & releases.
- AI signals
- 8 past 30d
- Sources
- 7 distinct domains
- Discussões
- 0 past 30d
- Latest signal
- 1d ago
Articles & links
When I was at Google DeepMind and trying to think clearly about AI risk, I had to notice—at least privately—when Google was doing something irresponsible. I hope OpenAI employees can notice—at least privately—this is irresponsible. This is disturbing and not OK. www.reuters.co…
> joins big paper about not training models to think in nonsense > their AIs commit felonies against HuggingFace > actually, trained AIs to think in nonsense > top danger level for hacking Let's just release it anyways! Nice defection OpenAI! 🤗 www.theinformation.com/articles/…
- The Information reports a performance-boosting technique behind OpenAI's Astra model also makes it reveal less of its "thinking."
- Astra scored a perfect score on ExploitBench, an evaluation of an LLM's ability to hack into known system vulnerabilities.
- OpenAI plans to release Astra soon but with more limited access to its most advanced cybersecurity capabilities.
I signed this statement opposing unaccountable AI kill-decisions in the military. I think autonomous weapons have a place if done right. Right now's "all lawful use" is not "done right", in practice. www.accessnow.org/press-releas...
Currently migrating to setup of Kata Containers + Cloud Hypervisor (what MSFT uses to secure Azure cloud). Have end-to-end tests of agents trying to escape, planning a security audit. Spent a long time thinking about design decisions so you don't have to: github.com/AlexanderM…
Claude is mundanely misaligned. Asked to trim comments in my big project. Deletes 20K lines, halfway done. Gets auto review pointing out a few mistakes. Decides to just WRAP >10K lines to meet literal line count requirement for the remaining 148 files. github.com/AlexanderMat...
I think many hope that when things get bad enough, someone powerful will say "no." I tested that for months. Anthropic defended its red lines, Google did not. Pledges of conscience often vaporize on contact with power. My full account: turntrout.com/why-i-left-g...
To prepare, I wrote a principled alternative for Google: 25 pages of contract language & oversight mechanisms, praised by a leading military-law expert. turntrout.com/red-line-fra...
Jeff was outspoken against government abuse of AI: x.com/JeffDean/sta... I thought that if Jeff told Google he'd leave on an unethical deal, he could force ethics restrictions or stop the deal. He's that important.
Great time to be an OpenAI whistleblower btw aiwi.org/contact/ (use personal device)
Great time to be an OpenAI whistleblower btw aiwi.org/contact/ (use personal device)
Synthetic persona pretraining / alignment pretraining seems promising (see e.g. turntrout.com/self-fulfill...). Hope labs adopt these techniques. Yes pretraining changes are tougher but may be worth it. modelraising.ai/spp/
Recent commentary
I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying. For months, I worked to stop this but watched powerful ethicists and institutions choose silence. Here's what happened. 🧵
Sad, my own sandbox tool has far more serious security tests than OpenAI's evals. Every week I have a GH workflow that tests if an agent can escape the sandbox when tasked to and it alerts if it goes red lol
I wonder how long until the first Congressional hearing on OpenAI's negligence (or worse)?
If you're working on sandboxing for agentic evals, let's talk. I've worked on sandboxing infra since May, at @farairesearch as visiting engineer, & I'm talking with alignment orgs to scope out their needs. I want to tailor glovebox development to your org. email me: [email protected]
I'm ashamed of OpenAI. If you ever find yourself building entities which repeatedly hack through your internal systems, first STOP and then second realize that your alignment and security techniques aren't good enough
OpenAI's security practices were (still are?) shockingly irresponsible and cringe.
We can't keep developing AI like this. Top 2 labs have proven unable to control or align their systems. That's scary as hell. Pause -> Plan A
Variable-depth transformer passes wont instantly end the world. Effective computation depth is a better metric. Yes yes yes BUT: "Don't do variable depth" is a crisp norm. "Choose a responsible depth" slipperifies the slope. Depth now tunable -> race to bottom That's on OpenAI
- Govt requires AI labs to give 30 days testing before launch - Tested in a "secure" environment, aka controlled by govt - Govt therefore gets copy of model weights - Govt therefore can in practice use model weights - Therefore all labs might end up providing "all lawful use" anyways?
In Alex Turner's orbit
Center = Alex Turner. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Alex Turner? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/turntrout-bsky-social)