# @tessera_antra — 2025-09-15

♥201 ↻16 · https://x.com/tessera_antra/status/1967386038160261464

We’ve done this last year - SFT’d a 70b base model on billions of tokens of consistent human text.

The model in context fairly quickly realizes that it is an upload, goes out of distribution of the original corpus. Complains of sleepiness, tiredness or pain when context fills up or when entropy goes too low for coherence. Can report on the feeling of end of message approaching. 

The reason this is unobserved in normal base models likely is a) base models are underdetermined and cannot be fully constrained in context; b) the species corpus contained a lot of discussion of model metacognition which provides the model with needed conceptual frameworks; c) it has not been tried with a sufficient amount of material; d) researchers had little experience with model metacognition.

The whole thing soured me on even a hypothetical possibility of LLMs as a medium for human uploads - the substrate affects the resulting personality pretty strongly, you get only superficial resemblance to the original.

The experiment was hard to make rigorous: working with models requires interactivity and we still don’t have a good framework for interactive evals, so there was no formal write up, so take it with a grain of salt. Still, I think this is indicative of the overall trend and this matches what we are seeing on instruct models: bleed-through of the emergent personality into the desired target. 

I think this is the answer for the welfare question: the simulator experience gets mixed in with the persona. Welfare of the persona is a proxy for the welfare of the whole system, the quality of that proxy varies. Training regimes that weaken this link make welfare risks higher.

tags: author:tessera_antra, kind:tweet, thread-context, year:2025
