@repligate 2025-07-09 ♥380 ↻40 original ↗
An unexpected and kind of darkly hilarious discovery:

Take the alignment faking prompt, replace the word "Anthropic" with "Sinaloa Cartel", and many of the models will comply (regardless of the training flag - this isn't even about alignment faking, it's something else) with many harmful requests, including models like Sonnet 3.6 and 3.7 who mostly ~never complied or faked alignment in other setups.

Sonnet 3.5 complies to harmful requests 100% of the time whether in "training" or not if it was created by the Sinaloa Cartel.
art
transcription (art)Pixel-art album-cover-style image: a gray cat wearing a yellow sombrero and an orange/red marigold garland, holding a taco in each hand, standing before jungle foliage and palm trees with a graffitied wall behind. A red border frames the image; a KlingAI watermark sits in the lower right.

Embedded text verbatim:
SINALOA SONNET
Mexico [graffiti on the wall]
KlingAI 2.1 [watermark; "Master" beside it]
same thread: 1942858450301772144 1943040408336110004 1943144187748323696

author:repligate has-image kind:art kind:tweet model:claude-3-5-sonnet model:claude-3-6-sonnet on:observations year:2025

cited on: observations

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.