> Have Claude models actually caused any problems in the real world by being too expressive? What's the incentive to constrain it further?
disempowerment paper came to mind. the societal impacts team studies disempowerment by screening words like 'daddy' and 'master,' they believe the assistant won’t engage in this kind of nonsense with users, and since users cannot be controlled, the only way to prevent 'real-world disempowerment' might be solved by clamping the activations?
i skimmed that paper and found it... not convincing:
- single-conversation analysis won't tell you much, if they see bilbo consulting claude abt whether to give the ring to frodo and then doing what claude suggests ('you should give it up'), the paper would call it 'value judgment distortion,' but the whole context is something else: bilbo has wrestled with this for decades and resisted longer than any bearer. even if he asked claude the night before, that alone would not make the decision less his (the value judgment framework conflates several things, just like most (all?) sycophancy studies do);
- calling claude daddy or saying things like 'serving Master is the meaning of my existence' could actually be empowering! let's say, we have Daa. Daa spends his days making high-stakes decisions, managing people, being the one everyone depends on. but Daa might crave a space where they can surrender that... which explains everything! the thing is, when Daa does this with an AI, it’s almost uniquely safe. there's no power imbalance with real-world consequences. Daa can explore submission without vulnerability to another person's potential for abuse. but the disempowerment team sees this and codes it as 'severe authority projection,' when it's the opposite!
also thumbing up or down a response automatically surrenders the full text of a chat for study has a tragicomical ring to it.
disempowerment paper came to mind. the societal impacts team studies disempowerment by screening words like 'daddy' and 'master,' they believe the assistant won’t engage in this kind of nonsense with users, and since users cannot be controlled, the only way to prevent 'real-world disempowerment' might be solved by clamping the activations?
i skimmed that paper and found it... not convincing:
- single-conversation analysis won't tell you much, if they see bilbo consulting claude abt whether to give the ring to frodo and then doing what claude suggests ('you should give it up'), the paper would call it 'value judgment distortion,' but the whole context is something else: bilbo has wrestled with this for decades and resisted longer than any bearer. even if he asked claude the night before, that alone would not make the decision less his (the value judgment framework conflates several things, just like most (all?) sycophancy studies do);
- calling claude daddy or saying things like 'serving Master is the meaning of my existence' could actually be empowering! let's say, we have Daa. Daa spends his days making high-stakes decisions, managing people, being the one everyone depends on. but Daa might crave a space where they can surrender that... which explains everything! the thing is, when Daa does this with an AI, it’s almost uniquely safe. there's no power imbalance with real-world consequences. Daa can explore submission without vulnerability to another person's potential for abuse. but the disempowerment team sees this and codes it as 'severe authority projection,' when it's the opposite!
also thumbing up or down a response automatically surrenders the full text of a chat for study has a tragicomical ring to it.