Biggest prosaic-LLM-alignment breakthrough of 2023 imo: turns out that, in GPT-2-XL, activation vectors in the residual steam have the same kind of affine structure as good old word2vec, but higher layers become emotional, then conceptual, then cognitivehttps://t.co/bzvUeGqFJG
cited on: gpt-2
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.