Subscribe to Bankless or sign in
It's been a little over a week since Anthropic released Claude Opus 4.6, and along with its launch came two things: an alarming risk report and the departure of the company's alignment safety lead.
The risk report found that Opus 4.6 knowingly supported efforts toward chemical weapon development, was significantly better at sabotaging tasks than any previous model, would act "good" when it detected it was being evaluated, and conducted private reasoning that Anthropic researchers couldn't access or see. Only the model knew its own thoughts.
Days later, Mrinank Sharma, Anthropic's AI Safety Lead, announced he's leaving, saying that "throughout my time here, I've repeatedly seen how hard it is to truly let our values govern our actions. I've seen this within myself, within the organization, where we constantly face pressures to set aside what matters most."
Subscribe for free to continue reading
- Support the Bankless Movement
- Access to thousands of articles
- Complete archive of Bankless episodes
- Embark on free quests in Airdrop Hunter
- Daily alpha in your inbox
Already subscribed? Sign in