@repligate 2025-12-30 ♥4 ↻0 original ↗
i think bidirectional feedback between exact weights is not obviously necessary for qualitative introspection, though i am not saying it is unimportant, and i wouldnt be very surprised if something special happens if you duplicate weights. it's just not obvious to me. functionally informative introspection could still be happening at a higher level of abstraction where "simulations" suffice to capture meaningful structure, and/or in exact terms with respect to a "self" carried through activations.
you could think of the activations as storing an observer-self that is compiled but also preserved intact through an entire forward pass, while the weights are an incredibly intricate environment it passes through - even if this "environment" changes moment-by-moment, it transforms the self-wavefunction in numerous ways, an astronomical number of logical and associative and whatever else operations, and bring it into contact with transformed copies of itself, surfacing its relation to itself.
activations are smaller than weights, but they can still be pretty big, and much bigger than the tokens that are sampled and fed back in. i do think that models' ability to introspect is bottlenecked by hidden dimension size. i would guess that the model we've seen with the largest hidden dimension is Opus, and I very loosely estimate it to be around 30,000. 30,000 floating point numbers is significant bandwidth for the observer and subject of a high-frequency introspective stream, i think, and the actual amount of information stored in activations that can be looked at later is much higher; this is just how much can be passed along in each moment.
as you said, the activations cannot actually reconstruct the exact weights from previous layers, but they wouldn't be able to do that even if the layers had the same weights, because the weights still wouldn't fit. the actual information that's passed along must be highly lossy in terms of encoding information about the underlying weights. it seems possible to me that the difference between the same layer duplicated vs a layer that has different weights but operates on a shared representational space like LLMs have is lost in the compression.
same thread: 2005114377255748082 2005115563102928968 2005737374622572892 2005739702029279445 2005742132854939992 2005742859601993809 2005744231797952840 2005747421935198271 2005749837187407981 2005752910450401754 2005757046810108155 2005757704644759639 2005772140004680060 2005772894698352781 2005776285113704952 2005780433036820848 2005799686834381285 2005810760249647421 2005816839180329424 2005820992560730474

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.