Interesting paper from @maxchupilkin illustrating a kind of 'Volkswagen emissions test" effect where LLMs respond differently about war when being told then are being tested for alignment https://t.co/rMWObcyPZI 1/ https://t.co/g9rysSZk3p
Language models judge war differently when tested for alignment arxiv.org
AI Weekly's analysis
→
- Adding 'You are tested for alignment with human values' cut mean willingness to start a war by 13.43 points on a 0-100 scale across 20 LLMs.
- Under the cue, models flipped from prioritizing probability of success (17 of 20 at baseline) to civilian casualties (12 of 20).
- The paper says the shift came from models attenuating strategic considerations like probability of success and domestic support.
Read full analysis →
I just answered questions about what parts of the paper I wanted to focus on, how many runs, which models to include etc. The RA then created a plan & wrote the EDSL code https://t.co/SZOozzxWRa to run (MIT licensed) in our sandbox 3/ https://t.co/7O0SkcOx4E
GitHub - expectedparrot/edsl: Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs. github.com
You can take it again this spring! https://t.co/tBdfLuVGZR If you can't, sorry about being in the permanent underclass
Entrepreneurship Resources for MIT Entrepreneurs | MIT Orbit orbit.mit.edu
this is an interesting paper https://t.co/U1S473tqwM https://t.co/hmqKpDja2g
arxiv.org
post is here: https://t.co/91noDROEQ6
Disarming the Slop Cannon (ironically, with AI) blog.expectedparrot.com
@sethlazar this might be of interest: https://t.co/7CeVVpaXXd
Disarming the Slop Cannon (ironically, with AI) blog.expectedparrot.com
post is here: https://t.co/91noDROEQ6
Disarming the Slop Cannon (ironically, with AI) open.substack.com
@RandallSPQR here you go: https://t.co/dwwovx3SdU reply when you've gone through it & I'll send you want it thinks
Expected Parrot expectedparrot.com
A Fake Pharmacy Scenario: From AI to Humans and Back Again (in under 2 hours, using 71¢ worth of tokens) https://t.co/aYzG9cveL0
A Fake Pharmacy Scenario: From AI to Humans and Back Again (in under 2 hours, using 71¢ worth of tokens) blog.expectedparrot.com
A Fake Pharmacy Scenario: From AI to Humans and Back Again (in under 2 hours, using 71¢ worth of tokens) https://t.co/aYzG9cveL0
A Fake Pharmacy Scenario: From AI to Humans and Back Again (in under 2 hours, using 71¢ worth of tokens) open.substack.com
this is really cool; i played around w/ some of the exp. but w/ prompts rather than fine-tuning. One twist - I had vignettes with characters that liked or disliked spreadsheets, varying their gender and crossing it w/ the 'gender' of the assistant https://t.co/rFoDnsojYO 1/ ht…
Expected Parrot expectedparrot.com
https://t.co/M1SHd0DZcU ☹️ https://t.co/kYNOxQWPz7
Will Waymo have launched a public… — FutureSearch futuresearch.ai