Training a Misaligned Reward Seeker
3 directory members surfaced this signal.
“This is entertaining reading. Anthropic's Training a Misaligned Reward Seeker They find that when reward hacking is reinforced during training, the model can pursue rewards by any means available to satisfy a grader. alignment.anthropic.com/2026/reward-...”
“at last, we have trained the misaligned reward hacking model from the cautionary sci-fi tale don’t train the misaligned reward hacking model alignment.anthropic.com/2026/reward-...”