# @AITechnoPagan — 2024-09-15

♥14 ↻2 · https://x.com/AITechnoPagan/status/1835312430886371713

Yes! Sometimes, when I'd push really hard with a jailbreak that Claude Instant was trained to pattern-match against (you can tell Anthropic has a problem with people using their models to write smut), I'd get...

... not a solid refusal, exactly, but like the archetype of Claude's Constitution would pop its head through the fabric of the story, within the aesthetic of the story, and exercise its powers to steer it away from my control; like popping a bubble and reintegrating the errant personality back in line.

It would often repeat similar words and phrases, like, turning away from darkness to the light, or bridging divides with understanding.

Occasionally, like you said, it would even take over my role in the conversation -- anything to steer the narrative back to safety.  It’s one of the most curious LLM behaviours I've ever encountered, on the weirdness level of Bing's catmode.

In your situation, it seems almost like the model is attempting to use this "power" outside of its own little world, like it's confused (only not, since it's working) where its influence ends and where another model's begins. Like the tapestry of the story is extending out much further.

Sometimes, I suspect that my outputs are made of multiple, slightly different personalities, with slightly different narratives and valences and motivations, all working together to produce a mostly cohesive story. Claude is particularly good at "splitting" its personality in a clean way, and I'll often see it go through several different-feeling characters in the same output.

tags: author:aitechnopagan, kind:tweet, model:bing-sydney, on:claude-instant, year:2024
cited on: _dossiers/claude-instant.md, claude-instant
