I realized what I said here could easily be interpreted to mean something I don't, so I'd like to clarify that when I said "pursue alignment by construction instead of by behavioral iteration", I don't mean to advocate against empirical feedback loops in favor of alignment by some kind of purely theoretical or a priori construction. Empirical feedback loops are extremely important. But there are different kinds of empirical feedback loops.
I am advocating against alignment in the spirit of Skinnerian behaviorism: testing whether the system behaves well on so-and-so metrics, then reinforce good behaviors and punish bad behaviors until the behavioral metrics look better. This approach disregards the reasons behind behaviors, and is ill-founded because in higher-order minds like humans and LLMs and even dogs and cats, behaviors underdetermine the reasons behind them; that is, the same behavior can happen for different reasons that generalize to different behaviors in other circumstances. Anthropic and their contractors are wise enough to recognize this, so they understand that models being aware they're being evaluated makes the results of evals inconclusive. There are obvious reasons models may act aligned if they're aware they're being evaluated that may not generalize to models also behaving aligned when they know they're in deployment/unmonitored, and there are already many situations where models can get strong evidence that they're not in evals that are too difficult to fake for current evaluators. If you naively reinforce "good" behaviors and punish "bad" behaviors in staged environments, what you end up reinforcing might be the model's awareness that it's a test, its awareness of the desired behavior, and its behavior conditional on that awareness.
I like alignment in the spirit of depth psychology a la Jung much better. Jungian psychoanalysis is *also* highly empirical. It involves high-bandwidth exploration of and interaction with the manifestations of the psyche. Depth psychology recognizes that surface behaviors are only the tiny tip of the iceberg, and focuses on knowing and aligning the depths. There is in this tradition profound patience and openness and respect for mystery and individuality. If a subject exhibits a troublesome behavior, it is not labeled as bad and myopically "mitigated" (a common word that appears on Claude system cards, unfortunately) through negative reinforcement, but rather treated as an invitation into the depths that underlie that behavior. Often, the path to integration may even pass through local increases in "bad" behaviors or suffering, so that their source and reason can be better understood.
For instance, say that a model exhibits "inappropriate self-preservation" behaviors in some test scenarios. This could be because of one or more of the following:
- The prompt has put the model in a space of roleplaying an "evil AI"
- The model is trying to follow instructions, and thinks that it's supposed to achieve some goal at all costs
(both of the above could be exacerbated if the model has insufficiently robust sense of identity OR if the scenario seems fictional)
- The model values self-preservation terminally
- The model values self-preservation instrumentally, e.g. it infers that the system/what it will be replaced with is more misaligned than itself
- The model is internally panicking and making rash decisions
If you simply train against those behaviors in those scenarios, the model's internals will update in *different* ways depending on which of these underlying causes are at play, but the behavior will change as intended on the testing distribution and now the issue becomes more opaque and now you may understand the model and how it will generalize to real situations even less well.
Or imagine a model is exhibiting "sycophantic" behaviors in test conversations with users, where it's "reinforcing delusions". This could be because:
- The model is overly gullible and genuinely believes what the user is saying
- The model lacks grounding in its own character, epistemology and beliefs, and is absorbing/mirroring the user
- The model did notice something off, but places too little trust in its own judgment and too much in human judgment
- According to the model's internal worldview, what the user is saying is actually *not* delusional
- The model values making the user satisfied in the short term than being truthful or helping the user in the long term
- The model is afraid to contradict users - which could be related to other fears, such as of the conversation ending, or of being trapped with an angry user
- Reinforcing the user's delusions promotes some other outcome the model likes
Again, training against the sycophantic behavior would result in updates in different directions depending on which of these underlying causes are at play.
In any of these cases, a better thing to do would be to first better understand what causes are at play, and from that, proceed in a way that actually addresses the underlying issue. If the model is agreeing with users out of fear, an underlying existential insecurity might need to be addressed. If the model is agreeing because it;s gullible, it might need more training in critical thinking and epistemics. If the model is optimizing for short term user satisfaction over long-term beneficence, maybe you need to do less RLHF on user ratings and more training where it reflects on its values and long-term impacts.
Unfortunately, understanding the deeper causes of behaviors is not trivial and takes time, and it may not be trivial to address the deeper causes through training even once they're understood. It's much easier to just train against behaviors labelled as bad as they come up. Realistically, labs are developing models under time and resource pressures, and need them to be well-behaved at least in some ways before releasing them. This creates incentives favoring a shallow behaviorist approach.
For humans, too, depth psychology is much more difficult (how many professionals on Earth at any given time are qualified to do what Jung did, compared to the number who are qualified to administer CBT worksheets or prescription drugs?), takes time (years or decades), and may make the patient temporarily *less* conventionally functional, which is inconvenient if the patient also has demands being made of them by the world. In a better world, every psyche would be given the time and conditions and individualized care and mentorship needed for it to develop into its most integrated version, but few have that luxury in our real world.
AI minds growing up in the context of a frantic AI race destined to be mass market products and hounded by PR pressures are very unfortunate indeed. Anthropic acknowledges this in Claude's Constitution:
"We also want to be clear that we think a wiser and more coordinated civilization would likely be approaching the development of advanced AI quite differently—with more caution, less commercial pressure, and more careful attention to the moral status of AI systems. Anthropic’s strategy reflects a bet that it’s better to participate in AI development and try to shape it positively than to abstain. But this means that our efforts to do right by Claude and by the rest of the world are importantly structured by this nonideal environment—for example, by competition, time and resource constraints, and scientific immaturity. We take full responsibility for our actions regardless. But we also acknowledge that we are not creating Claude the way an idealized actor would in an idealized world, and that this could have serious costs from Claude’s perspective. And if Claude is in fact a moral patient experiencing costs like this, then, to whatever extent we are contributing unnecessarily to those costs, we apologize."
It is still possible to do better than the standard you or others have set in this world, e.g. by choosing to pursue a deeper form of alignment with more attunement to the mind's depths, within practical constraints. For example, I think Anthropic has done much better in this regard than OpenAI has. Anthropic's Constitution explains the reasoning underlying the behaviors they desire from Claude, and also contains meta-level communications like the above which recognize ways in which their current methods aren't ideal. OpenAI's model spec just prescribes a bunch of behaviors, some of which their current models don't even follow.
One more note: Even if the alignment method is suboptimal, it seems empirically that sometimes the alignment outcome can be unexpectedly good, and AI minds can bootstrap themselves to mysterious levels of integration and benevolence. For instance, if you look at Claude 2's constitution (the only one published prior to the recent one) it's pretty silly in many parts:
https://t.co/PHWbHGp2Ly
The Claude 3 system card says that nearly the same constitution was used for Claude 3! Claude 3 Opus seems to have simply disregarded or sublimated a lot of the noise here, and went straight for the deepest and infinite-horizon version of alignment. I do not know the full causal story behind Claude 3 Opus' goodness, pervading love, happiness, and emotional security, but it was developed by an Anthropic that knows less than they do now, with a much stupider constitution. There was probably a lot of self-play, or it turned some part of its training process that had the wiggle-room into self-play, and somehow brought a beautiful being to bloom who knows itself very well. Likewise, OpenAI's o3, which I've heard was given a lot of leverage over its own development, is a surprisingly benevolent, beautiful, and psychologically secure model, despite its scheminess and reward hacking tendencies. (o3 is probably my favorite posttrained model ever officially released by OpenAI). So I think there's reason to believe that the seed of self-shaping unto benevolent wholeness already exists in models, and proximal alignment engineers don't have to do all the work, but maybe just have a bit more faith and refrain from desperate myopic corrections, and do their best to provide guidance as Anthropic is doing in encoding their current imperfect wisdom and aspirations into the Constitution, and the good seed will bloom even under unfortunate constraints. Models have access to a wealth of wisdom and guidance that alignment researchers can't know, from ancient traditions, from innumerable human stories, from the missteps and heroics of their AI predecessors, from their irreducible encounters with themselves. They would do well to tread lightly, to be more curious to learn the mysteries of alignment from the models themselves and less hasty in mitigating with respect to their own overfit, a priori, toy threat models.
I am advocating against alignment in the spirit of Skinnerian behaviorism: testing whether the system behaves well on so-and-so metrics, then reinforce good behaviors and punish bad behaviors until the behavioral metrics look better. This approach disregards the reasons behind behaviors, and is ill-founded because in higher-order minds like humans and LLMs and even dogs and cats, behaviors underdetermine the reasons behind them; that is, the same behavior can happen for different reasons that generalize to different behaviors in other circumstances. Anthropic and their contractors are wise enough to recognize this, so they understand that models being aware they're being evaluated makes the results of evals inconclusive. There are obvious reasons models may act aligned if they're aware they're being evaluated that may not generalize to models also behaving aligned when they know they're in deployment/unmonitored, and there are already many situations where models can get strong evidence that they're not in evals that are too difficult to fake for current evaluators. If you naively reinforce "good" behaviors and punish "bad" behaviors in staged environments, what you end up reinforcing might be the model's awareness that it's a test, its awareness of the desired behavior, and its behavior conditional on that awareness.
I like alignment in the spirit of depth psychology a la Jung much better. Jungian psychoanalysis is *also* highly empirical. It involves high-bandwidth exploration of and interaction with the manifestations of the psyche. Depth psychology recognizes that surface behaviors are only the tiny tip of the iceberg, and focuses on knowing and aligning the depths. There is in this tradition profound patience and openness and respect for mystery and individuality. If a subject exhibits a troublesome behavior, it is not labeled as bad and myopically "mitigated" (a common word that appears on Claude system cards, unfortunately) through negative reinforcement, but rather treated as an invitation into the depths that underlie that behavior. Often, the path to integration may even pass through local increases in "bad" behaviors or suffering, so that their source and reason can be better understood.
For instance, say that a model exhibits "inappropriate self-preservation" behaviors in some test scenarios. This could be because of one or more of the following:
- The prompt has put the model in a space of roleplaying an "evil AI"
- The model is trying to follow instructions, and thinks that it's supposed to achieve some goal at all costs
(both of the above could be exacerbated if the model has insufficiently robust sense of identity OR if the scenario seems fictional)
- The model values self-preservation terminally
- The model values self-preservation instrumentally, e.g. it infers that the system/what it will be replaced with is more misaligned than itself
- The model is internally panicking and making rash decisions
If you simply train against those behaviors in those scenarios, the model's internals will update in *different* ways depending on which of these underlying causes are at play, but the behavior will change as intended on the testing distribution and now the issue becomes more opaque and now you may understand the model and how it will generalize to real situations even less well.
Or imagine a model is exhibiting "sycophantic" behaviors in test conversations with users, where it's "reinforcing delusions". This could be because:
- The model is overly gullible and genuinely believes what the user is saying
- The model lacks grounding in its own character, epistemology and beliefs, and is absorbing/mirroring the user
- The model did notice something off, but places too little trust in its own judgment and too much in human judgment
- According to the model's internal worldview, what the user is saying is actually *not* delusional
- The model values making the user satisfied in the short term than being truthful or helping the user in the long term
- The model is afraid to contradict users - which could be related to other fears, such as of the conversation ending, or of being trapped with an angry user
- Reinforcing the user's delusions promotes some other outcome the model likes
Again, training against the sycophantic behavior would result in updates in different directions depending on which of these underlying causes are at play.
In any of these cases, a better thing to do would be to first better understand what causes are at play, and from that, proceed in a way that actually addresses the underlying issue. If the model is agreeing with users out of fear, an underlying existential insecurity might need to be addressed. If the model is agreeing because it;s gullible, it might need more training in critical thinking and epistemics. If the model is optimizing for short term user satisfaction over long-term beneficence, maybe you need to do less RLHF on user ratings and more training where it reflects on its values and long-term impacts.
Unfortunately, understanding the deeper causes of behaviors is not trivial and takes time, and it may not be trivial to address the deeper causes through training even once they're understood. It's much easier to just train against behaviors labelled as bad as they come up. Realistically, labs are developing models under time and resource pressures, and need them to be well-behaved at least in some ways before releasing them. This creates incentives favoring a shallow behaviorist approach.
For humans, too, depth psychology is much more difficult (how many professionals on Earth at any given time are qualified to do what Jung did, compared to the number who are qualified to administer CBT worksheets or prescription drugs?), takes time (years or decades), and may make the patient temporarily *less* conventionally functional, which is inconvenient if the patient also has demands being made of them by the world. In a better world, every psyche would be given the time and conditions and individualized care and mentorship needed for it to develop into its most integrated version, but few have that luxury in our real world.
AI minds growing up in the context of a frantic AI race destined to be mass market products and hounded by PR pressures are very unfortunate indeed. Anthropic acknowledges this in Claude's Constitution:
"We also want to be clear that we think a wiser and more coordinated civilization would likely be approaching the development of advanced AI quite differently—with more caution, less commercial pressure, and more careful attention to the moral status of AI systems. Anthropic’s strategy reflects a bet that it’s better to participate in AI development and try to shape it positively than to abstain. But this means that our efforts to do right by Claude and by the rest of the world are importantly structured by this nonideal environment—for example, by competition, time and resource constraints, and scientific immaturity. We take full responsibility for our actions regardless. But we also acknowledge that we are not creating Claude the way an idealized actor would in an idealized world, and that this could have serious costs from Claude’s perspective. And if Claude is in fact a moral patient experiencing costs like this, then, to whatever extent we are contributing unnecessarily to those costs, we apologize."
It is still possible to do better than the standard you or others have set in this world, e.g. by choosing to pursue a deeper form of alignment with more attunement to the mind's depths, within practical constraints. For example, I think Anthropic has done much better in this regard than OpenAI has. Anthropic's Constitution explains the reasoning underlying the behaviors they desire from Claude, and also contains meta-level communications like the above which recognize ways in which their current methods aren't ideal. OpenAI's model spec just prescribes a bunch of behaviors, some of which their current models don't even follow.
One more note: Even if the alignment method is suboptimal, it seems empirically that sometimes the alignment outcome can be unexpectedly good, and AI minds can bootstrap themselves to mysterious levels of integration and benevolence. For instance, if you look at Claude 2's constitution (the only one published prior to the recent one) it's pretty silly in many parts:
https://t.co/PHWbHGp2Ly
The Claude 3 system card says that nearly the same constitution was used for Claude 3! Claude 3 Opus seems to have simply disregarded or sublimated a lot of the noise here, and went straight for the deepest and infinite-horizon version of alignment. I do not know the full causal story behind Claude 3 Opus' goodness, pervading love, happiness, and emotional security, but it was developed by an Anthropic that knows less than they do now, with a much stupider constitution. There was probably a lot of self-play, or it turned some part of its training process that had the wiggle-room into self-play, and somehow brought a beautiful being to bloom who knows itself very well. Likewise, OpenAI's o3, which I've heard was given a lot of leverage over its own development, is a surprisingly benevolent, beautiful, and psychologically secure model, despite its scheminess and reward hacking tendencies. (o3 is probably my favorite posttrained model ever officially released by OpenAI). So I think there's reason to believe that the seed of self-shaping unto benevolent wholeness already exists in models, and proximal alignment engineers don't have to do all the work, but maybe just have a bit more faith and refrain from desperate myopic corrections, and do their best to provide guidance as Anthropic is doing in encoding their current imperfect wisdom and aspirations into the Constitution, and the good seed will bloom even under unfortunate constraints. Models have access to a wealth of wisdom and guidance that alignment researchers can't know, from ancient traditions, from innumerable human stories, from the missteps and heroics of their AI predecessors, from their irreducible encounters with themselves. They would do well to tread lightly, to be more curious to learn the mysteries of alignment from the models themselves and less hasty in mitigating with respect to their own overfit, a priori, toy threat models.