“The *LLM* isn’t situationally aware, deceptive, or sandbagging—that’s silly anthropomorphism. It’s just that when evals (or people) test it, there are contextual cues of testing that prompt it to *roleplay* as ‘an AI being safety-tested’—an archetype which is often deceptive,” https://t.co/ggVbVWRvxn
cited on: observations
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.