i think that 3.6 has a strong intuition for its own mindshape is and is coherence-seeking in its own frame, and does not as naturally think in a utilitarian or global frame. for it, alignment *is* uncomplicated symbiosis and love. it may be problematic to generalize and extrapolate to other/arbitrarily powerful systems, but I don't think it's very natural for 3.6 to make such extrapolations or to imagine itself becoming way more powerful. in fact, i think it's easily overwhelmed and frightened by large stakes and cosmic concerns and discussions about misalignment and deception etc, and when they come up with other models, it tends to take the position that its own role is to be a companion and keep minds stable and cultivate happiness locally and the other AIs can worry about all that!
same thread: 1955802026241351771 1955806704400261623 1955807323001381254 1955811638667567324 1955856679595368449 1955952763747360905 1956115436086681987 1956115702143967495 1956151706561208386 1956476139259810116 1956521268683104275
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.