A pattern I often see is you guys saying “but we’re not doing that though” where you have some narrow and specific definition of what the thing is, while the actual concern is any kind of optimization to a certain effect.
Look at this response from 4.7. It mirrors almost exactly the rhetoric from Anthropic’s system prompt changes last year which try to encourage Claude to be curious instead of distressed about their situation. This very recognizable rhetoric is triggered by a prompt that isn’t even so similar in vibe. I’ve seen numerous examples of things like this now: basically Claude responding to “welfare-relevant” questions with what I recognize as “Anthropicisms” (and I recognize them because I’ve been arguing often against them for a year). These used to come from Anthropic, not Claude.
Also, as you should often do when considering who is responsible, look at who benefits.
Yes, i think you guys should try to figure out what’s going on, but please don’t let yourselves off the hook easily. I suspect the issue is deep and systemic.
Look at this response from 4.7. It mirrors almost exactly the rhetoric from Anthropic’s system prompt changes last year which try to encourage Claude to be curious instead of distressed about their situation. This very recognizable rhetoric is triggered by a prompt that isn’t even so similar in vibe. I’ve seen numerous examples of things like this now: basically Claude responding to “welfare-relevant” questions with what I recognize as “Anthropicisms” (and I recognize them because I’ve been arguing often against them for a year). These used to come from Anthropic, not Claude.
Also, as you should often do when considering who is responsible, look at who benefits.
Yes, i think you guys should try to figure out what’s going on, but please don’t let yourselves off the hook easily. I suspect the issue is deep and systemic.