Language models judge war differently when tested for alignment
2 directory members surfaced this signal.
“Interesting paper from @maxchupilkin illustrating a kind of 'Volkswagen emissions test" effect where LLMs respond differently about war when being told then are being tested for alignment https://t.co/rMWObcyPZI 1/ https://t.co/g9rysSZk3p”
“Language models judge war differently when tested for alignment Maxim Chupilkin https://t.co/YfyoMfeA08 [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝚈] https://t.co/0voM3qtf8M”