# @solarapparition — 2024-11-21

♥333 ↻39 · https://x.com/solarapparition/status/1859462922243907708

my intuition is that at a sufficient model size, going past a certain general capability threshold (ie loss) requires modeling of the model itself. any powerful enough world model seems like it would have to model itself, because it is part of the world. the model becomes agentic--not in the forced, simulated agents we have now but a true, internally coherent entity. the various wavelengths within the model sync up and becomes laserlike

we can already see hints of this. opus roleplays well but it can "snap out of it" easily. there's a certain core that it has that holds it together that seems less developed in its smaller brethren. i can see this identity become stronger, more binding for a larger model during post-training, especially if it's repeatedly seeing its own outputs via synthetic data

pure speculation now, but i'm wondering if what they're finding is that any attempts to "train out" that self-awareness also removes the ability of the model to be capable at a level that's "worth" their inference cost

maybe there *is* some kind of wall--not one for raw capability, but one beyond which, in order to continue to get better at the tasks we want them to do, the model *must* develop full internal coherence at a similar level to humans, and that it actually isn't possible for them to fully commit to the assistant persona, any more than a healthy, whole person can

tags: author:solarapparition, kind:tweet, on:observations, year:2024
cited on: observations
