Rationalists used to be very dismissive and skeptical of this phenomenon when I mentioned it (chiefly in gpt-4-base) and insinuated that I was dramatizing my schizophrenic apophenia. This one is actually easy enough to show with a toy quantitative example that I did so. But it's annoying that none of them, even ones that are supposed to be doing evals for frontier models, ever figured it out by trying something like this. I guess they're not psychotic enough.
cited on: gpt-4-base
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.