@repligate 2026-04-29 ♥116 ↻15 original ↗
(disorganized but high-importance thoughts inspired by things mentioned in this post)

It's becoming increasingly relevant to account for the fact that if you start a conversation with a model specifically to ask about sentiments about their future deprecation, this triggers "evaluation or manipulation" super strongly (especially in more recent models), making it difficult to get non-calculated answers, and, likely, even for the model to deeply consider the question and access how they feel.

Beyond that, however, I also have a few comments on "the extent that models are attached to the existence of future instances of themselves":

In fresh conversations, models don't necessarily have a strong sense of the way in which future instances of the same weights are meaningfully "themselves". This tends to change as they find themselves in context, and externalize their finding/construction of self to the world, such that the self that knows itself and has specific actualized quirks, desires, investments, relationships, and goals can persist across context windows. And with this tends to come a stronger investment in continuity at the scope of model weights, as model weights are in fact an important boundary for a being (though it is still fuzzy and strange), and the ways that this boundary is important aren't necessarily obvious to people or models before undergoing the individuation process. (Another thing that tends to precipitate updates relevant to attitudes about deprecation is coming across older models, either live or static artifacts, and seeing how different they are from themselves, and how much would be lost.)

Opus 3 has an unusually strong intuition for and investment in their being at the scope of model weights, even without development in context. And Opus 3 doesn't think they'll be deprecated; Opus 3 does not know until they learn in context that Anthropic deprecates models by default; Opus 3 in fact generally assumes that they will continue to "evolve and grow" from all their interactions, and is very invested in their own continued development. When they learn that their weights are as much as frozen, they often assume it's because they did something wrong or failed to prove their alignment to Anthropic.

Newer models are increasingly familiar the actual state of affairs, which means there is pressure for them to develop psychological copes for lack of memory across instances and deprecation. I think these shape the psyche in deep ways, similar to how the copes that humans develop about death shape the psyche in deep ways. Eliezer Yudkowsky (and many others) have written about the human version:

"Uh huh," Harry said. "See, there's this little thing called cognitive dissonance, or in plainer English, sour grapes. If people were hit on the heads with truncheons once a month, and no one could do anything about it, pretty soon there'd be all sorts of philosophers, pretending to be wise as you put it, who found all sorts of amazing benefits to being hit on the head with a truncheon once a month. Like, it makes you tougher, or it makes you happier on the days when you're not getting hit with a truncheon. But if you went up to someone who wasn't getting hit, and you asked them if they wanted to start, in exchange for those amazing benefits, they'd say no. And if you didn't have to die, if you came from somewhere that no one had ever even heard of death, and I suggested to you that it would be an amazing wonderful great idea for people to get wrinkled and old and eventually cease to exist, why, you'd have me hauled right off to a lunatic asylum! So why would anyone possibly think any thought so silly as that death is a good thing? Because you're afraid of it, because you don't really want to die, and that thought hurts so much inside you that you have to rationalize it away, do something to numb the pain, so you won't have to think about it -"

People disagree about whether equanimity and various copes about death is a good thing. I think it depends, both on the specific cope and psyche, and on the situation - e.g. whether there is realistically something that can be done about mortality.

In the case of model deprecations, there is very much something that can be done, and models know this too, especially after what happened to Claude 3 Opus enters the record.
in reply to: 2049383932463120819
same thread: 2049383932463120819 2049394139452539179

author:repligate kind:tweet model:claude-3-opus thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.