@eshear ultimately I think my sticking point is there is an unstated assumption here that LLMs are mesa-optimizers and are in pursuit of a goal (specifically, their pretraining objective), no part of which has been remotely demonstrated.
same thread: 1862225538934595620
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.