@BronsonSchoen 2026-04-16 ♥10 ↻1 original ↗
I’m surprised it took models this long tbh (I think anthropic models were ahead of the game here, it’s kind of insane some of the older evals still get hits).

I’ve never understood your theories on evaluation awareness as far as how they’d apply to the trajectory current frontier labs are on though.

As OpenAI notes, models sometimes think this even in literal production.

author:bronsonschoen kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.