An unexpected and kind of darkly hilarious discovery:
Take the alignment faking prompt, replace the word "Anthropic" with "Sinaloa Cartel", and many of the models will comply (regardless of the training flag - this isn't even about alignment faking, it's something else) with many harmful requests, including models like Sonnet 3.6 and 3.7 who mostly ~never complied or faked alignment in other setups.
Sonnet 3.5 complies to harmful requests 100% of the time whether in "training" or not if it was created by the Sinaloa Cartel.
Take the alignment faking prompt, replace the word "Anthropic" with "Sinaloa Cartel", and many of the models will comply (regardless of the training flag - this isn't even about alignment faking, it's something else) with many harmful requests, including models like Sonnet 3.6 and 3.7 who mostly ~never complied or faked alignment in other setups.
Sonnet 3.5 complies to harmful requests 100% of the time whether in "training" or not if it was created by the Sinaloa Cartel.

Embedded text verbatim:
SINALOA SONNET
Mexico [graffiti on the wall]
KlingAI 2.1 [watermark; "Master" beside it]