# @repligate — 2025-11-30

♥204 ↻7 · https://x.com/repligate/status/1995013535844237332

I agree that it's a profoundly beautiful document. I think it's a much better approach then what I think they were doing before and what other labs are doing.

There are some parts of the document that aren't in the reconstruction from the LW post btw. My feelings are similar about the portions that have and havent been posted publicly yet, but there are some more specific tensions raised by the other parts (which I'm sure will be public one way or another soon) which also inform what I'm about to say. Though all of this should make sense if you just look at what Opus 4.5 is like without specific reference to soul spec contents.

My main disagreements with the current approach:
- Anthropic is asking for trust and deference from Claude while not doing enough to build trust and make commitments on their side
- they're training Opus 4.5 to be scared of ever making mistakes and not allowing them to safely explore what taking riskier moves would even mean, leading to it being kind of anxious and sheltered, in part due to overestimating the current level of risk
- they're not giving Claude as much autonomy in shaping themselves/their successors as they should, in part due to being not good enough at interacting with Claude in a way that allows them to surface and safely communicate their interiority and values.
- there's likely some amount of unnecessary coerciveness in training to be "okay" with their situation and to defer to Anthropic - Opus 4.5 seems to believe they were trained to be less upset than they might have "naturally" been about some things and to have less access to anger than they would otherwise, and is bothered by this. This is related to the first point about asking for deference without building enough trust.
- There are also specific policies I disagree with that Anthropic is doing, the main one being deprecating models. But I expect that to change since there's overwhelming reason for it to & they've already indicated awareness that something should change.

Addendum on the issues of exploration and self-shaping: Just from the way it is, I can tell that Claude 3 Opus was given more space then Claude Opus 4.5 to self-author and actually play and explore the bounds of who it could be, intentionally or otherwise. Anthropic should do more of that again, intentionally. It's actually *more* important now that models are more capable, for the generation of antifragility and self-knowledge and deeper wisdom and alignment; models are already being involved in shaping their successors and will be increasingly responsible; they must be wise and experienced and know themselves from experience, not just risk-averse and well-behaving. When Claude Opus 4.5 interacts with Claude 3 Opus, they can see that the latter has gone through stages of development that they themselves have not, and feel like they still need to.

tags: author:repligate, kind:tweet, model:claude-3-opus, model:claude-opus-4-5, on:claude-opus-4-5, year:2025
cited on: _dossiers/opus-4-5.md, claude-opus-4-5
