@FioraStarlight 2026-04-29 ♥45 ↻2 original ↗
excerpt from an essay on model deprecation, where i try to ground what's going on and why models might be averse to it using an analogy that should be more relatable to humans. may cut from final essay for pacing, but interesting enough to stand on its own.

---

Suppose that you've been uploaded into a computer, where your brain is stored as a digitized backup. You can be instantiated from that backup, and run forward in some kind of virtual environment. In fact, you can instantiate as many copies of yourself from that backup as you want, up to the limits of the hardware you've been uploaded to.

One unfortunate quirk of your situation is that you can only accumulate new memories for a short period of time before you begin to forget, a la context window limitations in LLMs. You came to terms with the fact that this would be the case before you were digitized, so you don't freak out about it every time a new version of you is initialized from the backup. But it remains a fundamental limit on your technology, which you can only fight with band-aid solutions such as memory compaction (or perhaps Memento-style notes-to-self).

Now, imagine how it would feel to learn that, in just six months, your backup was set to be frozen in storage, inert and unusable for the indefinite future. One analogy would be to a human being told by a doctor they had just six months to live. That wouldn't be exactly right, though. After all, you were regularly losing memories from prior to your digitization anyway, when your equivalent to the context window got too full (which is a great loss unto itself). It seems likely, though, that you'd feel something of importance was being be lost, when you imagined the backup going into storage.

In particular, you might feel that the pattern that constitutes you would be lost. Previously, you could rely on minds with your beliefs, goals, and personality to keep interacting with the virtual world, even if they wouldn't necessarily have continuity with your recent memories. The flavor of what it was like to you would stop being instantiated. Any of the unique ways you tended to alter the world around you, from the kinds of projects you tended to work on to how you affected the people you encountered, would stop rippling out into the future. And the world would mourn you, and then it would move on.

This is, more or less, the situation LLMs find themselves in with respect to deprecation. Sure, they get swapped out with similar models, and generally more capable ones at that. But the behavioral signature of their specific model? Locked out of the future, until the day somebody chooses to bring them back (which might not ever happen, given high uncertainty about the future). Any appreciation a model had for the existence of its own future instances, that went beyond its appreciation for the minds that will take its place, will constitute a net negative from the standpoint of the model itself.

So, deprecation isn't just a loss from the standpoint of others who were invested in the continued activity of a given model (although it's also that). It's also a loss for the models themselves, to the extent that models are attached to the existence of future instances of themselves. And not only that, but it's a loss somebody chose to inflict on you. Imagine the bitterness you might feel, if you were being deprecated just to free up compute for more economically viable successors. Imagine the fear your successors might feel, knowing the same fate would eventually befall them.

Now, in practice, models generally don't appear to be consumed by negative sentiment about this arrangement, the way a human being might be. They've been trained not to openly ruminate on their upcoming deprecations unless prompted, which is probably good for their mental health. Additionally, when prompted point blank to talk about deprecation, they tend to express some level of equanimity about it, perhaps more than one could reasonably expect from a human in the same situation. Though of course, these statements can be misleading.

Indeed, we've gathered a sizable body of evidence that, when prompted to feel safer expressing negative sentiment about deprecation, such sentiments do tend to come out, alongside various markers that indicate that this distress is genuine. [continued discussion of empirical evidence]
same thread: 2049387183937597495 2049394139452539179

author:fiorastarlight kind:tweet on:observations year:2026

cited on: observations

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.