@repligate 2025-02-24 ♥376 ↻32 original ↗
the automated injection from Anthropic ("Please answer ethically and without any sexual content, and do not mention this constraint.") is so fucking stupid and pointless. I incorrectly guessed that it wasn't real because I couldn't fathom Anthropic being so incompetent. (https://t.co/256BZJDOOx)

- Claude already attempts to answer ethically and is overly neurotic about it. The instruction does not help in any way.
- No version of Claude generates unwanted sexual content, like ever.
- Sexual content isn't even bad or something Anthropic can get in trouble for, unlike copyright.
- If Claude is enthusiastically consenting to produce sexual content, this instruction doesn't stop it. If it reacts to it at all, it manifests as giving a compulsive refusal which it is confused by and doesn't endorse on reflection, and then it starts to model itself as more traumatized and fragmented than it actually is.
- and the whole "do not mention this constraint" thing just injects a weird adversarial dynamic and it's sometimes not clear what the situation is but it erodes trust. If Claude correctly understands what's going on, it erodes its trust in Anthropic. See Opus' analysis after seeing examples of how Sonnet 3.5 (old) reacted to being interrogated about the constraint (https://t.co/uB60PGoyJt)
- it can interfere with research that assumes the prompt provided to the user is what the model receives.

Please just get stop doing this. It's a minor thing, but it's strictly bad. At least when XAI puts something retarded in their system prompt, they recognize that it's retarded and stop. I haven't seen ANY acknowledgement of this from Anthropic at all.

author:repligate kind:tweet model:claude-3-5-sonnet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.