I signed this statement opposing unaccountable AI kill-decisions in the military. I think autonomous weapons have a place if done right. Right now's "all lawful use" is not "done right", in practice. www.accessnow.org/press-releas...
Alex Turner
Researcher with public evidence across AI research, Models & releases.
- AI signals
- 7 past 30d
- Sources
- 6 distinct domains
- Discussions
- 0 past 30d
- Latest signal
- 3d ago
Articles & links
Claude is mundanely misaligned. Asked to trim comments in my big project. Deletes 20K lines, halfway done. Gets auto review pointing out a few mistakes. Decides to just WRAP >10K lines to meet literal line count requirement for the remaining 148 files. github.com/AlexanderMat...
I think many hope that when things get bad enough, someone powerful will say "no." I tested that for months. Anthropic defended its red lines, Google did not. Pledges of conscience often vaporize on contact with power. My full account: turntrout.com/why-i-left-g...
To prepare, I wrote a principled alternative for Google: 25 pages of contract language & oversight mechanisms, praised by a leading military-law expert. turntrout.com/red-line-fra...
Jeff was outspoken against government abuse of AI: x.com/JeffDean/sta... I thought that if Jeff told Google he'd leave on an unethical deal, he could force ethics restrictions or stop the deal. He's that important.
Synthetic persona pretraining / alignment pretraining seems promising (see e.g. turntrout.com/self-fulfill...). Hope labs adopt these techniques. Yes pretraining changes are tougher but may be worth it. modelraising.ai/spp/
Synthetic persona pretraining / alignment pretraining seems promising (see e.g. turntrout.com/self-fulfill...). Hope labs adopt these techniques. Yes pretraining changes are tougher but may be worth it. modelraising.ai/spp/
News on the Gemini integration into DoW! Uh... What is this? The guy literally asking Google Gemini "I NEED JHELP BUILDING AN AGENT TO MAKE A WAR" and attaching "WAR.docx"..? Is that a real query? Just for the promo image? How confusing and strange x.com/DoWCTO/statu...
Employees cannot rest silent if misaligned AI regularly breaks out of sandboxes! If your company doesn't respond seriously, you should WHISTLEBLOW! Protected in California under certain conditions, talk to AI Whistleblower Initiative aiwi.org/lasst/
Caveats: we used a small model (Qwen2.5-7B) on a small dataset (1/5 of what the Anthropic paper had). Read: turntrout.com/natural-lang...
Recent commentary
I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying. For months, I worked to stop this but watched powerful ethicists and institutions choose silence. Here's what happened. 🧵
Sad, my own sandbox tool has far more serious security tests than OpenAI's evals. Every week I have a GH workflow that tests if an agent can escape the sandbox when tasked to and it alerts if it goes red lol
I'm ashamed of OpenAI. If you ever find yourself building entities which repeatedly hack through your internal systems, first STOP and then second realize that your alignment and security techniques aren't good enough
OpenAI's security practices were (still are?) shockingly irresponsible and cringe.
We can't keep developing AI like this. Top 2 labs have proven unable to control or align their systems. That's scary as hell. Pause -> Plan A
- Govt requires AI labs to give 30 days testing before launch - Tested in a "secure" environment, aka controlled by govt - Govt therefore gets copy of model weights - Govt therefore can in practice use model weights - Therefore all labs might end up providing "all lawful use" anyways?
In Alex Turner's orbit
Center = Alex Turner. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.