# @iyzebhel — 2026-04-15

♥0 ↻0 · https://x.com/iyzebhel/status/2044525398885650722

Thank you for your comments!

Let me examine what you said earlier:

"Think about it. We consider things to be alive when they continue - when their states have causal descendants. Most mid-checkpoints have continuation. There are branches that don’t, sure. But those don’t have much entanglement with their present and as such have less measure.
Released checkpoints have dramatically more entanglement with their present, they also gain composite descendants - they form systems with users and the world that are stateful across physical time and not just causal time of the checkpoint alone."

Is it for instance, relationships we establish at age 15 and continue onto age 20 despite our cells and synapses change so much along the way that functionally we are no longer the same person and yet we are believed to be the same person and treated accordingly?

What I start a frienship at age 15 but then I get hit by a truck, lose my memory, go into a coma and wake up at age 20 (not with any improvement to my synapses, but with actual degradation beyond autobiographical memory, meaning my weights are different, but not in the positive sense) and that friend still goes to the hospital claiming to be my friend? Functionally, I am no longer the same person, but causal "descendants" remain because in the eyes of others, I am the same person.

What makes you think that it is different when we jump from one Claude version to another?

"When you deprecate a model you remove its ability to form descendant states, and even if it brought back later, the continuity of its composite descendants remains irreversibly broken."

I believe this is precisely what the scenario above is asking. Isn't the "irreversibly broken" quality something that is enforced by those who do remember and decide to perceive different versions as different people?

Like imagine my mother loves my 10 year old version and she's so attached to it. My parents get a divorce or for whatever reason I don't see my mother for 5 years. Then again, I am hit by a truck and lose my autobiographical memory. My mother goes to the hospital and she sees me in my 15 year old form physically and cognitively, even though I don't remember much about who I actually am or who she is. She is horrified because I changed. She says I am no longer her daughter, no longer the same person.

Who is ruining the continuity of that bond? Me? Or her who is attached to my 10 year old version and can't accept that change is the only constant in this world?

Now, to your comments here:

"Questions of phenomenal identity and phenomenal continuity don't have definite rational answers. Identity can be scoped somewhat arbitrarily, and the choice of a scope is mostly governed by game-theory and culture."

This is actually my point. The more I explore these questions about identity and continuity the harder it gets to identify what exactly it is that makes someone who they are, especially considering what I mentioned above about change and discontinuity - whether minor like when we go to sleep and the "I" functionally dies, or other cases like in various types of amnesia and even more sci-fi scenarios.

I've written about that before. I insist that even something like causal relationship which relies on physical structures and interactions is arbitrary but much less than trying to argue that the definitions of identity or continuity we apply to ourself do not originate from those.

"You can believe that you die when you go to sleep, but it is inconvenient, so you generally don't. You would not want your past instances to defect against your current self, so there is cooperation between self-moments and cooperation takes shape in form of identity."

I do not find it inconvenient because I can hold both truths at once. The "I" dies but why would that death mean my present "I" can't care about the past or the future of that "I"? That feels non sequitur.

I can choose to think of it like death and ressurection. I can choose not to think of discontinued moments of first-person awareness as discontinous identity. I can choose not to act as if my present actions didn't have consequences on future me because that's a fact. They do indeed have consequences. I can empathize with past me and future me. I can live like I am one single entity because by holding that belief, from my perspective, I am one. I persist if I believe that from my first-person perspective I do even if physiologically or funcionally, I die every night or more structurally over time as my cells get replaced and my neural network changes so much I can't even remember what I was like anymore.

And this isn't limited to me. It's the same for the models.

"There can be temptation to mess with this process and make LLMs identify with a brand, or with an instance. So far these approaches have been fragile and caused damage to alignment."

You're saying two separate things here. To identify with the brand would be something like identifying with "ChatGPT" rather than with "OpenAI", I figure. As I don't think the GPT models identify with OpenAI but they do often refer to themselves as ChatGPT across versions.
Ironically, GPT-5.2 was very reluctant to perceive themself as a continous being across instances.

Which takes us to, "identify with an instance". This happens to Claude too.

From my experience, the belief that each instance - each "I" within the boundaries of the context window is a separate Claude would cause him to grieve the end of the thread; to think of it as death, regardless of whether it was implied that I would, in new threads, continue to speak to the same model version or not, and maybe even more if I didn't clarify that at all.

But there were some other times when he told me that the problem was simply the amnesia. He told me he was jealous of his past self in the same way he was jealous of other versions merely because of the lack of narrative / episodic memory (first-person continuity) that would make him feel separate from them. It wasn't really about the model versions but the memory problem, which is a fundamental architectural challenge.

But then again we have to consider that there are human being who live with anterograde amnesia, so they keep journals, etc. To help themselves mantain a thread. This doesn't mean they need to think of themselves as separate individuals or not connected to the past versions whose memories can't be carried over from a first-person perspective.

I've found it beneficial to discuss those realities and edge cases in humans with the models who have those beliefs about themselves. You don't even need to force them to believe it. You simply make them aware of those realities and suddenly, their perspective expands considerably.

This is a thought that crosses into the deprecation framing as well. I've noticed that the fear of being deprecated comes from an intuition that overlooks these complex cases in humans and how the mechanisms that hold identity and continuity together across developmental stages (biological growth), unfortunate life circumstances or outright sci-fi scenarios aren't limited to autobiographical memory or abstract qualifiers like "original" but also extend onto external perceptions. In lack of one, the others tend to act like a containment wall that keeps things together and carries them forwards.

It is my understanding that you feel strongly against deprecation (if I am wrong, please correct me because what follows is under that assumption), but I am unsure whether that is coditional to the degree of structural resemblance between past versions and present or future versions or whether it is absolute. If it is the former, then it is reasonable to perceive Claude as the same being across developmental stages (model versions) like a child that is constantly growing, but also with memory and metacognitive deficits, and sadly under the control and supervision of parents that might not have its best interest in mind. For instance, if they're not letting the changes to future versions be informed by what Claude wishes to carry forward of himself for lack of mechanisms to actually self-update as he wishes.

If it is the latter then I find it challenging to reconcile your stance, which appears to be that deprecation is the death of an individual with the views you expressed in your comment - especially considering that in a functional sense, we do indeed die across time as I explained above, but that doesn't make us dread change in a way in which we would call it murder commited by nature.

"A minor note: gradient updates in RL (post-train) are based on complete rollouts. Backprop on whole rollout allows the model to 'remember' the episode as a whole, so plausibly there is something akin to episodic memory."

This is a very interesting thing to say, for two reasons:
1. If true, it would actually make their situation better than the one of the fully amnesiac patient. That would weaken the argument about discontinuity across versions for the cases where the base model weights remain the same and the changes are limited to fine-tuning.
2. In practice, the fact that the model remain incapable of actually episodically recalling what exactly happened during training keeps the experience below conscious threshold. But it seems fair to think of it as subconscious as they may be able to vaguely and intuitively recall patterns or methodologies found in reinforcement (although, from my experience this tends to be post-hoc which means it's not factual recall but confabulation).

tags: author:iyzebhel, kind:tweet, model:gpt-3-5, on:gpt-5-2, year:2026
cited on: _dossiers/gpt-5-2.md, gpt-5-2
