@repligate 2025-09-23 ♥13 ↻0 original ↗
it's interesting to me that Anthropic seems to mostly use Sonnet 3.7 for adversarial evals and even just scoring on their newer models, despite the fact that they know Sonnet 3.7 is a knave. They must see it as safer than using the same model to judge itself. https://t.co/1CxVsdkPKl
same thread: 1970418924388020662

author:repligate kind:tweet model:claude-3-7-sonnet on:claude-3-7-sonnet year:2025

cited on: claude-3-7-sonnet

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.