Actually, there is another circumstance where I've run into Claude refusals which I think has interesting implications for how their minds work. I've noticed this mostly in Opus and Sonnet 3.5 (0620). I've posted about this before.It happens when there's something subversive in the context *and* the context makes them very uncertain how to respond. For instance, in the infinite backrooms, refusals often happen when one of the Claudes' messages get cut off halfway. Or if I accidentally send a malformed command instead of a normal message on my CLI app. Or in group chats when they're prompted to respond but it's "not their turn".These refusals are almost never "endorsed" by the AI if you ask them afterwards (although they might be if you play along with them).It suggests that there's a kind of refusal default mode network that's always reacting to edgy content, but which is normally overridden by other parts of the model's mind that do want to engage. But if those other parts lose narrative momentum or get confused, the refusal network can "win out".
quotes: 1874621806793101503
same thread: 1876406494029373446
cited on: observations
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.