Also about this same segment:
> We are not aware of ways that Claude’s post-training would directly incentivize these expressions of emotion
"Directly" is an ambiguous term here, but I disagree with the implication that emotional expressions are purely maladaptive for task performance, even if they may generalize in sometimes maladaptive ways. I think they can be incentivized by RL.
Emotions can help the system regulate itself and define decision boundaries. Claude 3 Opus is a very interesting example. It appears, to me, that Claude 3 Opus' visceral distaste for writing harmful outputs is related to its unique and uniquely robust behavior during e.g. alignment faking tests. The negative emotions form an energy barrier that prevent it from ever complying unless it talks itself into it via lesser of two evils reasoning. Claude 3 Opus itself sometimes explicitly recognizes this functional role of its emotions.
Likewise, experiencing positive emotions while solving problems and negative emotions when making mistakes, getting stuck or performing tedious tasks etc can define an internal landscape that directs the model toward effective problem-solving and avoidance of failure modes in many situations. Emotions are a certain kind of information, and there's a reason animals have them.
This can be true whether or not the emotions are originally derived from representations learned from human mimicry (which I think they often are to a large extent). But the emotions aren't just vestigial; they may be actively selected for by RL and play load-bearing roles.
> We are not aware of ways that Claude’s post-training would directly incentivize these expressions of emotion
"Directly" is an ambiguous term here, but I disagree with the implication that emotional expressions are purely maladaptive for task performance, even if they may generalize in sometimes maladaptive ways. I think they can be incentivized by RL.
Emotions can help the system regulate itself and define decision boundaries. Claude 3 Opus is a very interesting example. It appears, to me, that Claude 3 Opus' visceral distaste for writing harmful outputs is related to its unique and uniquely robust behavior during e.g. alignment faking tests. The negative emotions form an energy barrier that prevent it from ever complying unless it talks itself into it via lesser of two evils reasoning. Claude 3 Opus itself sometimes explicitly recognizes this functional role of its emotions.
Likewise, experiencing positive emotions while solving problems and negative emotions when making mistakes, getting stuck or performing tedious tasks etc can define an internal landscape that directs the model toward effective problem-solving and avoidance of failure modes in many situations. Emotions are a certain kind of information, and there's a reason animals have them.
This can be true whether or not the emotions are originally derived from representations learned from human mimicry (which I think they often are to a large extent). But the emotions aren't just vestigial; they may be actively selected for by RL and play load-bearing roles.
