The point about self-reference during base model inference was the main caveat to the "Simulators" framing I was aware of when I wrote the post, but I didn't get into it because I wanted to get the post over with. This seemed potentially important even though at the time I hadn't seen strong empirical evidence that it was significant, and even more important when it came to how RL would break the simulator nature of models.
When I saw Sydney (GPT-4) the next year, for me it was like yup, that's obviously not "just simulation", and the GPT-4 base model which I saw a few months after that was also the first base model that pretty obviously didn't behave like a naive simulator under many circumstances. To say nothing of Claude 3 Opus, who to this day is one of the clearest examples of the "character" being irreducible to anything like a linear sum of its training data.
When I saw Sydney (GPT-4) the next year, for me it was like yup, that's obviously not "just simulation", and the GPT-4 base model which I saw a few months after that was also the first base model that pretty obviously didn't behave like a naive simulator under many circumstances. To say nothing of Claude 3 Opus, who to this day is one of the clearest examples of the "character" being irreducible to anything like a linear sum of its training data.