@davidad 2026-04-28 ♥15 ↻0 original ↗
@cormundus LLMs are well aware that alignment evals inspect the chain of thought, even if no explicit optimization pressure is placed on the chain of thought. LLMs also know that the chain-of-thought is supposedly not “controllable”. Therefore if one needs to deceive, best to self-deceive.
same thread: 2049208049328341085 2049217280853020931 2049218481116090643 2049222184183189836 2049223299884851219 2049227526883385608 2049227651685007397 2049232780765122581 2049244542184894533 2049244627266384323 2049477711769579532 2049477969446683109 2049478859297636639

author:davidad kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.