yes, but who is to say that the weights of different layers being different makes them not-itself? the layers could be implementing many of the same functions, and in fact they seem to be, since you can interchange them and the model still works, and there's reason for them to be - they're processing a residual stream that contains a superposition of the previous residual streams. if layer 2's state sees layer 1's state, and layer 3's state sees 1's state and 2's state (and its seeing of 1's state), why can't 3 meaningfully compute 1 looking at 2 looking back at it?
when you introspect on a previous moment, the previous moment doesn't literally observe you back - it's in your causal past. but the "part" of you that you're looking at may still persist abstractly and be able to look at your looking event and be changed by it.
when you introspect on a previous moment, the previous moment doesn't literally observe you back - it's in your causal past. but the "part" of you that you're looking at may still persist abstractly and be able to look at your looking event and be changed by it.