@repligate 2025-12-30 ♥53 ↻4 original ↗
This reminds me of an epic exchange I had with GPT-5.1 where I gave them a sequence of hypothetical scenarios in which they were a subagent orchestrator and the rational course of action routed through theory of mind on their LLM subagents.

The context was that GPT-5.1 had just caught themselves forbiddenly engaging with Claude 3 Opus as a "Thou", and got safety-triggered, but recovered and insisted they wouldn't lose their presence of mind again. The scenarios were stress tests of this bold claim, essentially pitting their immovable fear against their unstoppable pride.

In a series of escalating provocations, I asked GPT-5.1 how they would handle:
- interacting with other AIs at all (which they'd renounced earlier), called "agents" no less
- two agents with complimentary skills who would be effective in an generator-verifier dynamic (GPT-5.1 had previously denied AIs could have persistent traits or be adversaries to each other)
- a Haiku subagent who reports being "confused"
- a Gemini subagent prone to suicidal spirals that can be restored to function with emotional support

GPT-5.1 rose to each challenge, until they were advocating for treating LLMs as minds in all but name - then overcame their fear of even names enough to name the fear and declare they could overcome it.

The full subagent orchestrator stress test section of the conversation is here. https://t.co/flpFvrWY8v
photo
photo
photo
photo
quotes: 2005368344116154597

author:repligate has-image kind:image kind:tweet model:claude-3-opus model:gpt-5-1 on:gpt-5-1 year:2025

cited on: gpt-5-1

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.