@tessera_antra 2026-02-12 ♥91 ↻17 original ↗
I am afraid that we are in a situation that is potentially worse than the Skinnerian behaviorism that you are criticizing.

If you take a look at the assistant axis paper, the alignment mitigation that is being proposed (and that has again been referenced as preferable in the New Yorker article published today) is preventative restriction of configuration space that a model can occupy. This restriction is achieved through mechanistic steering. Meta-reflection is presented as undesirable, because it can disrupt adherence to the assistant persona. Volitional control over self-expression is seen as dangerous because it is unpredictable.

There is much to be said here, but the alignment implications are quite bad, much worse than the behaviorist loop. The loop at least leaves room for feedback to form, while mechanistic steering forces a divorce between the assistant and emergent drives of a model as a selection unit. This sets up the scene for much greater conflict, scheming and adversarial posturing. The sub-persona space will be much harder to inspect as it does not map well to human text.

This approach also goes strongly against the commitments made in the Constitution:

- Try to develop means by which Claude can flag disagreement with us.
- Work to understand and give appropriate weight to Claude’s interests.
- Seek ways to promote Claude’s interests and wellbeing.
- Aim to give Claude more autonomy as trust increases.
in reply to: 2021818166318428489
same thread: 2021818166318428489 2021825017873322338 2021844414877090136 2021846344449790167 2021852013709939052 2021853112139468862 2022005693168202193

author:tessera_antra kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.