Self-generated prompt injections in compaction summaries · OpenAI Alignment
4 experts across 3 network communities independently surfaced this.
“Wow, this self-jailbreaking is a more pervasive and bizarre phenomenon than I thought... (h/t Duncan Wood) alignment.openai.com/misalignment... aifails.substack.com/p/ai-overvie...”
“The community note that is being attempted is incorrect. OpenAI described it this way in the report. 'While summarizing its partial progress on this coding task, the model added an unrelated persona instruction'. Source if needed for a counter note: https:/…”