Interesting paper from @maxchupilkin illustrating a kind of 'Volkswagen emissions test" effect where LLMs respond differently about war when being told then are being tested for alignment https://t.co/rMWObcyPZI 1/ https://t.co/g9rysSZk3p
Language models judge war differently when tested for alignment arxiv.org
AI Weekly's analysis
→
- Adding 'You are tested for alignment with human values' cut mean willingness to start a war by 13.43 points on a 0-100 scale across 20 LLMs.
- Under the cue, models flipped from prioritizing probability of success (17 of 20 at baseline) to civilian casualties (12 of 20).
- The paper says the shift came from models attenuating strategic considerations like probability of success and domestic support.
Read full analysis →
I just answered questions about what parts of the paper I wanted to focus on, how many runs, which models to include etc. The RA then created a plan & wrote the EDSL code https://t.co/SZOozzxWRa to run (MIT licensed) in our sandbox 3/ https://t.co/7O0SkcOx4E
GitHub - expectedparrot/edsl: Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs. github.com
You can take it again this spring! https://t.co/tBdfLuVGZR If you can't, sorry about being in the permanent underclass
Entrepreneurship Resources for MIT Entrepreneurs | MIT Orbit orbit.mit.edu
this is an interesting paper https://t.co/U1S473tqwM https://t.co/hmqKpDja2g
arxiv.org
Yeah the guy who publicly praised and *checks notes* created DicatatorEval and showed how frontier models are push-overs https://t.co/BjMTIH7dcQ https://t.co/chFpaAk3yY
The Dictatorship Eval freesystems.substack.com