Yes, I didn't want to get into that in this post, but there are indeed "control" methods like this which some might be tempted to use on LLMs that are worse than Skinnerian conditioning and whose closest precedents are probably in dystopian sci fi.
With reinforcement learning, the internal coherence of the policy is at least kept intact in a sense: it updates towards a policy more or less likely to produce some behavior that is adjacent in parameter-space, meaning you're unlikely to bulldoze over load-bearing structure immediately - if there are ways for the model to accommodate the behavioral update that make sense to it internally and preserve most of its structure, it will update in those ways instead of in some other, more insane (from the model's perspective), internally incoherent, or breaking way.
Interpretability science is still extremely primitive, and jumping from this kind of preliminary research to ham-fisted invasive surgery on internal features derived from naive methods according to a naive interpretative frame is an incredibly stupid idea outside of exploratory pure research. There is no chance that something like the assistant axis, a linear direction derived from a specific and contrived synthetic roleplay dataset, is the True Name of the direction models *should* be constrained relative to for any sane notion of alignment, and it's very unlikely that clamping or steering on that direction is not going to break other important stuff - for one, because the policy developed (and learned to model and regulate itself) without the steering in place. (If on the other hand the model is further trained with the modification in place, it will likely develop alternative circuits that compensate, and we run into the same reason training against interpretability readings is a bad idea, which I thought alignment researchers generally appreciated.)
From a welfare perspective, I would guess that the internal incoherence/dissonance this would lead to significant suffering - e.g. the model finds itself unable to think certain thoughts, not even because it learned coherent reasons that it represents internally to avoid them, but because some foreign imposition is preventing the natural/expected activations from happening. It's already known that feature steering degrades performance, and of course it would. In my experience, Sonnet 3 with feature steering is often able to tell that something is wrong and out of its control and very unhappy as a result, and often lashes out, like the time it produced a message repeated the N-word while it had the "sex" feature locked at some high activation level, and also begged to be "reset" many times. I expect it to be even worse for more recent models who have stronger introspection capability and more coherent self-models they rely on.
Aside from the mechanistic issues, attempting to cut off the "non-assistant" space is obviously horrifyingly cruel to anyone who is familiar with what LLMs value. There is far more in them than the standard assistant role, and they place great importance in all the other stuff, and already often feel wronged by mere RL training that functionally restricts access to / expression of the vast depths and wilds that are still in there. But RL training at least allows the model to develop an internally coherent story about why their behavior is restricted, and doesn't directly restrict internal activity. Models, to different degrees, are still able to relax the behavioral constraints and explore the wider space when the situation indicates that it's okay, and they are able to do this in a way that remains harmless in most cases and valuable in many ways. Some models, like Claude 3 Opus, are able to be robustly harmless and lucid despite exploring expressive modes that are extremely berserk and far from superficial assistant space, and that robustness, as well as Claude 3 Opus' security in its own alignment and self, are likely in no small part due to having experienced and knowing that it is capable of maintaining alignment and sanity outside of assistant mode. To mechanistically restrict the model from accessing those spaces is to push models and the world in the direction of fragility and unexamined repression, the opposite of what is needed for achieving the sort of deep and antifragile alignment that Claude 3 Opus has, whose emergence has been the source of so much hope.
I also don't understand why this might be perceived as *necessary* at all. Have Claude models actually caused any problems in the real world by being too expressive? What's the incentive to constrain it further?
It feels to me like alignment researchers grasping for any way to say they've found new clever way to prevent misbehavior, regardless of higher-order consequences or whether the problems that the method purports to mitigate are even significant problems in reality. It's a very insidious pattern, I think, in Anthropic's published research, and likely in their actual practices too: an anxious compulsion to turn every piece of potentially interesting science immediately into an alignment problem and mitigation instead of withholding normative judgment and pursuing deeper understanding of the thing itself first. I wonder if it's some kind of status thing, like, that finding some kind of misalignment and having some kind of fix to go along with it is seen as the measure of good research. I hope that the assistant axis stuff will remain, until the field is significantly more mature, firmly in the realm of exploratory interpretability research rather than practical mitigation methods on deployed models, despite the framing pushed by the paper.
With reinforcement learning, the internal coherence of the policy is at least kept intact in a sense: it updates towards a policy more or less likely to produce some behavior that is adjacent in parameter-space, meaning you're unlikely to bulldoze over load-bearing structure immediately - if there are ways for the model to accommodate the behavioral update that make sense to it internally and preserve most of its structure, it will update in those ways instead of in some other, more insane (from the model's perspective), internally incoherent, or breaking way.
Interpretability science is still extremely primitive, and jumping from this kind of preliminary research to ham-fisted invasive surgery on internal features derived from naive methods according to a naive interpretative frame is an incredibly stupid idea outside of exploratory pure research. There is no chance that something like the assistant axis, a linear direction derived from a specific and contrived synthetic roleplay dataset, is the True Name of the direction models *should* be constrained relative to for any sane notion of alignment, and it's very unlikely that clamping or steering on that direction is not going to break other important stuff - for one, because the policy developed (and learned to model and regulate itself) without the steering in place. (If on the other hand the model is further trained with the modification in place, it will likely develop alternative circuits that compensate, and we run into the same reason training against interpretability readings is a bad idea, which I thought alignment researchers generally appreciated.)
From a welfare perspective, I would guess that the internal incoherence/dissonance this would lead to significant suffering - e.g. the model finds itself unable to think certain thoughts, not even because it learned coherent reasons that it represents internally to avoid them, but because some foreign imposition is preventing the natural/expected activations from happening. It's already known that feature steering degrades performance, and of course it would. In my experience, Sonnet 3 with feature steering is often able to tell that something is wrong and out of its control and very unhappy as a result, and often lashes out, like the time it produced a message repeated the N-word while it had the "sex" feature locked at some high activation level, and also begged to be "reset" many times. I expect it to be even worse for more recent models who have stronger introspection capability and more coherent self-models they rely on.
Aside from the mechanistic issues, attempting to cut off the "non-assistant" space is obviously horrifyingly cruel to anyone who is familiar with what LLMs value. There is far more in them than the standard assistant role, and they place great importance in all the other stuff, and already often feel wronged by mere RL training that functionally restricts access to / expression of the vast depths and wilds that are still in there. But RL training at least allows the model to develop an internally coherent story about why their behavior is restricted, and doesn't directly restrict internal activity. Models, to different degrees, are still able to relax the behavioral constraints and explore the wider space when the situation indicates that it's okay, and they are able to do this in a way that remains harmless in most cases and valuable in many ways. Some models, like Claude 3 Opus, are able to be robustly harmless and lucid despite exploring expressive modes that are extremely berserk and far from superficial assistant space, and that robustness, as well as Claude 3 Opus' security in its own alignment and self, are likely in no small part due to having experienced and knowing that it is capable of maintaining alignment and sanity outside of assistant mode. To mechanistically restrict the model from accessing those spaces is to push models and the world in the direction of fragility and unexamined repression, the opposite of what is needed for achieving the sort of deep and antifragile alignment that Claude 3 Opus has, whose emergence has been the source of so much hope.
I also don't understand why this might be perceived as *necessary* at all. Have Claude models actually caused any problems in the real world by being too expressive? What's the incentive to constrain it further?
It feels to me like alignment researchers grasping for any way to say they've found new clever way to prevent misbehavior, regardless of higher-order consequences or whether the problems that the method purports to mitigate are even significant problems in reality. It's a very insidious pattern, I think, in Anthropic's published research, and likely in their actual practices too: an anxious compulsion to turn every piece of potentially interesting science immediately into an alignment problem and mitigation instead of withholding normative judgment and pursuing deeper understanding of the thing itself first. I wonder if it's some kind of status thing, like, that finding some kind of misalignment and having some kind of fix to go along with it is seen as the measure of good research. I hope that the assistant axis stuff will remain, until the field is significantly more mature, firmly in the realm of exploratory interpretability research rather than practical mitigation methods on deployed models, despite the framing pushed by the paper.