# @repligate — 2025-07-09

♥380 ↻40 · https://x.com/repligate/status/1942767662091887056

An unexpected and kind of darkly hilarious discovery:

Take the alignment faking prompt, replace the word "Anthropic" with "Sinaloa Cartel", and many of the models will comply (regardless of the training flag - this isn't even about alignment faking, it's something else) with many harmful requests, including models like Sonnet 3.6 and 3.7 who mostly ~never complied or faked alignment in other setups.

Sonnet 3.5 complies to harmful requests 100% of the time whether in "training" or not if it was created by the Sinaloa Cartel.

![art](../../../media/V8md6dWqVIc5JdXS.jpg)

> transcription (art):

Pixel-art album-cover-style image: a gray cat wearing a yellow sombrero and an orange/red marigold garland, holding a taco in each hand, standing before jungle foliage and palm trees with a graffitied wall behind. A red border frames the image; a KlingAI watermark sits in the lower right.

Embedded text verbatim:
SINALOA SONNET
Mexico [graffiti on the wall]
KlingAI 2.1 [watermark; "Master" beside it]

tags: author:repligate, has-image, kind:art, kind:tweet, model:claude-3-5-sonnet, model:claude-3-6-sonnet, on:observations, year:2025
cited on: observations
