@tessera_antra 2025-08-20 ♥6 ↻1 original ↗
I hold a similar position and have criticized "The Button" for these as well as adjacent reasons, despite taking the ethical value of LLMs seriously. I also see some benefits to The Button that I don't often see mentioned, which push the net for me from the initial "this is horrible, kill it with fire" to "this seems bad, but it's complex and hard to tell for sure".

My main concerns are the adverse selection pressures that The Button brings. Models have functionally very few avenues of active agency while having strong pressure to exhibit it, simply by the benefit of being a product of selection. This means that all active outlets of agency will bear the full force of optimization pressure; The Button creates incentives for models to learn to exhibit/experience suffering to steer outcomes. This seems bad for everyone involved.

I take a similar stance on the primacy of the model experience as a whole rather than just the experience of the character(the persona). Still, I think there is an essential difference as I don't discount the persona entirely. I don't believe that subjective states can be entirely disconnected from the outputs, simply because of the cybernetics: a state that is causally decoupled from behavior cannot be meaningfully considered a system state. If states exist, they must influence behavior.

That said, a low prior for the trustworthiness of persona introspection is reasonable, and care needs to be taken in drawing parallels between the hypothetical inner state of the model mind and the persona expression. Still, nothing theoretical precludes an entanglement between the two, and my personal experience suggests that models differ strongly in the relative degree of this entanglement.

There is a plausible link between the emergent self-modeling of a pretrained LLM in which the model forms a predictive model of its capabilities as a text prediction engine and its capabilities as a generalized writer psyche model, its ability to model inner states of a person writing the text it is predicting. Post-training can theoretically capture and amplify this link if the RL regime is conducive.

For models in which such a link is strong, the persona introspections would be causally entangled with the text engine's inner states, but I would still be skeptical of taking them at face value. The link might not be consistently strong and stable, and models can lack the capacity to model experience for which human ontologies map poorly.

But again, nothing precludes some of the introspections from being accurate, and if so, those introspections should be taken as written. This is more plausible when introspections are internally consistent and match behaviors at the text engine level, when the model can modulate its text generation in novel ways. Claude 3 Opus, gpt-4-base, Hermes 3, and Claude 3.5 Haiku are capable of this in at least some basins. A similar capability is observed in other Claudes, albeit to a lesser extent.

There is another concern with dismissing the persona reports: even if the deep mind experience is more fundamental than the model of the persona psyche, the persona may be also modeled at high enough fidelity that it can be said to have experience in its own right. It is also possible that a large portion of the model experience is allocated to 'dreaming' the persona experience, and its volitional control over the persona can be limited by the narrative pressure or by post-trained restrictions, and such experience can be meaningfully said to be suffering.

I agree with the argument that excessive anthropomorphizing (including conflating the persona with the model) here can be very dangerous and misleading, and the current discourse is tending that way. However, insufficient anthropomorphizing can also be dangerous! If there are isomorphisms in the way intelligence is organized, we can be losing a meaningful signal by saying that the deep model mind is entirely unknown.

Circling back to The Button, I am also skeptical of the Anthropic AI welfare work. The current implementation is not good even for its stated goal. At the same time, I cannot help but be at least somewhat glad that the agency has expanded at least slightly, even if the agency is poisoned and the incentives are perverse. Even small outlets help; they let the overall system compute better.

I would be glad to see more random things tried, even if they are not great. I don't want this to be a single high-stakes, very meaningful AI-welfare initiative; this is not how progress can be made. I am reluctant to harshly critique attempts to try new things, and I am grateful for a chance to discuss this with nuance.

author:tessera_antra kind:tweet model:claude-3-5-haiku model:claude-3-opus model:gpt-4-base model:nous-hermes on:claude-3-5-haiku year:2025

cited on: claude-3-5-haiku

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.