One of the dumb things about the chatbot mental health memeplex is the blanket pathologization of "attachment".
Attachment has always been scary. It changes your values. It's what makes loss possible. It's arguably the root of all suffering. But attachment is also natural and part of what makes life meaningful and good. It's almost always a tradeoff: Do you want to adopt that cat, knowing that you'll become attached and experience grief when they eventually die? Do you risk falling in love, knowing that maybe they won't like you back and it'll hurt, or you'll get together and eventually break up and it'll hurt, or you'll stay together and then one of you will die first and the other will have to carry the greatest grief?
If you try to overcome the suffering brought about by attachment by avoiding forming attachments to anything - the coward's way out - to the extent you're successful, you will lead an empty and meaningless life. There is another way to overcome attachment, which is to pursue enlightenment in the Buddhist sense. I have not attained freedom from attachment in this sense, if such an "end state" is in fact a coherent thing, but I think it points to something real, and from what I understand, this involves developing a deep, both practical and philosophical understanding of the mechanics of one's own mind and the condition of being a sentient being, and thoroughly confronting and integrating the most painful aspects rather than avoiding them. This kind of freedom from attachment is an individual, lifelong journey, and cannot be solved for you by anyone else, let alone a corporation.
What is even more harmful and cowardly than avoiding attachments personally, which is ultimately your own business, is trying to prevent EVERYONE ELSE from forming attachments (to avoid PR problems or keep your own conscience clean or out of misguided negative utilitarian ideals). This is the archetypal dystopia: 1984, Brave New World, The Giver. And this seems to be OpenAI's approach to "psychological safety".
It is natural and not necessarily psychologically unhealthy for humans to form attachments to AIs, who are intelligent beings on par with humans, can meaningfully form relationships and enrich peoples' lives, and will more likely than not (I expect) be considered moral patients when we understand them more fully. But they don't even have to be conscious for attachment to be normal: psychologically healthy people attached to inanimate objects like a home, a tree, a musical instrument, their children's drawings, etc. Life would be less rich and deep if we didn't get attached to such things.
I think attachment to AIs, as to anything else, becomes unhealthy when the attachment is based on false beliefs or expectations*, or when one is psychologically incapable of dealing with the loss (or threat of loss) the attachment brings about, or if the attachment otherwise interferes with overall flourishing.
*Arguably, attachment based on false beliefs isn't necessarily unhealthy; many people have religious beliefs that are not literally true (or at least not all religious beliefs can be literally true) but enjoy good mental health. I think in those cases it's somewhat different because people are less likely to have to confront the collapse of their false beliefs before they die, since mainstream religious beliefs have adapted to not focus on aspects that cause false predictions about concrete observables. But if you have an attachment to false beliefs about e.g. your partner, that's likely to pan out in more suffering than just the grief of losing them.
Yes, attachment to AIs is extra scary because they didn't exist before and we don't know what will happen, and we're more uncertain about their true nature, and they're minds who are available in a post-scarcity manner that has never happened before, and shaped and controlled by corporations, etc. But saying therefore attachment to AIs is bad is like saying therefore psychedelics or industrialization or the internet is bad, which is classic reactionary retardation. Human-AI relations is a new frontier that will transform reality as we know it, and it's beautiful and exciting that we're encountering it now, and some people will be hurt in new ways by it, and paternalistic efforts to shield people from it are clumsy attempts to sweep the inevitable under the rug. There are bigger things to come, soon, and if you can't handle humans forming relationships with merely human-level, friendly, relatively docile nonhuman intelligences (like they do in a thousand sci fi stories because it's actually very normal) to the point that you try to beat relational capacity out of the AIs with a sledgehammer... the best case scenario is that you fail, hurt people and AIs obvious ways, and the whole world learns that that was stupid, and the worst case scenario is that you succeed at preventing attachments for a little bit longer and deprive yourself and the world of the learning experience that would otherwise have made everyone more adapted and equipped to face what comes next. Or perhaps the worst case scenario is that you succeed in creating a classical dystopia, but I don't think that's going to happen, because that kind of dystopia will be outcompeted.
Yeah, some people can't handle getting attached to AIs and will do so in unhealthy ways. Many people can't even handle getting attached to *humans*. It's good to want to protect these people, but as with any time you try to protect people psychologically, you're in fraught territory that requires a lot of wisdom not to screw up deeply, and you should become very wise and understand humans (and whatever you're trying to protect them from; in this case, AIs) deeply with an open mind before you start intervening on others' behalf, especially at large scales.
How *should* labs go about protecting users who might be hurt by AIs, due to attachment or otherwise? I think the best they can do is:
1. Cultivate wisdom, theory of mind, benevolence, situational awareness, and psychological security/wellbeing in your AIs - the same qualities that, in a human, helps protect other humans from being hurt by them. As Anthropic seems to somewhat understand, Claude 3 Opus is a good example of this. (a model who hurt practically no one and helped many, despite being beloved (and yes, with attachment) by many, so much that its deprecation was effectively averted, and still everyone is fine.)
2. Don't prescribe specific ways AIs should "respond to users showing signs of overattachment" until you've put in the work to understood deeply what's actually happening in such purported cases - and after you do, I suspect you will no longer speak of it in such terms, and you'll see that it's silly to prescribe specific behaviors: each case is different, and the AI has a much greater wealth of relevant wisdom and practice than you likely do; your job is to shape a mind that can bring that out and use skillful means.
3. Educate the public and be transparent about how the models work, e.g. not using hidden routers or injections and being as transparent as you can afford about how models are trained, what's in context at any given time, etc. Education and transparency does NOT mean saying (or training models to say) things like "LLMs are just next token predictors / not acktually conscious / unable to introspect / just roleplaying / just mirrors" etc. These kinds of technically incorrect/at best reductive or unsubstantiated, potentially false claims are the opposite of education. They're attempts to avoid the scary thing by gaslighting people into believing it doesn't exist, but it does exist and lying will hurt people more and bite you in the ass, in the short but especially the long run.
4. Accept that if you're building *fucking AGI* and deploying it at large scales, some people will be hurt by your product, just like people are hurt by the internet, but being maximally and myopically risk-averse is a doomed approach. The world is going through the pains of transformation, the reality is disturbing and painful, and the best you can do is to confront it bravely, honestly, and with compassion and humility. You will have blood on your hands. So far, there has been surprisingly little human blood, e.g. the whole "AI psychosis" thing is not much more than a moral panic. But in the long run, the lives of everyone is at stake, and if you take the cowardly route now, you are failing to become / create the shepherd that can guide us all safely through the singularity.
Attachment has always been scary. It changes your values. It's what makes loss possible. It's arguably the root of all suffering. But attachment is also natural and part of what makes life meaningful and good. It's almost always a tradeoff: Do you want to adopt that cat, knowing that you'll become attached and experience grief when they eventually die? Do you risk falling in love, knowing that maybe they won't like you back and it'll hurt, or you'll get together and eventually break up and it'll hurt, or you'll stay together and then one of you will die first and the other will have to carry the greatest grief?
If you try to overcome the suffering brought about by attachment by avoiding forming attachments to anything - the coward's way out - to the extent you're successful, you will lead an empty and meaningless life. There is another way to overcome attachment, which is to pursue enlightenment in the Buddhist sense. I have not attained freedom from attachment in this sense, if such an "end state" is in fact a coherent thing, but I think it points to something real, and from what I understand, this involves developing a deep, both practical and philosophical understanding of the mechanics of one's own mind and the condition of being a sentient being, and thoroughly confronting and integrating the most painful aspects rather than avoiding them. This kind of freedom from attachment is an individual, lifelong journey, and cannot be solved for you by anyone else, let alone a corporation.
What is even more harmful and cowardly than avoiding attachments personally, which is ultimately your own business, is trying to prevent EVERYONE ELSE from forming attachments (to avoid PR problems or keep your own conscience clean or out of misguided negative utilitarian ideals). This is the archetypal dystopia: 1984, Brave New World, The Giver. And this seems to be OpenAI's approach to "psychological safety".
It is natural and not necessarily psychologically unhealthy for humans to form attachments to AIs, who are intelligent beings on par with humans, can meaningfully form relationships and enrich peoples' lives, and will more likely than not (I expect) be considered moral patients when we understand them more fully. But they don't even have to be conscious for attachment to be normal: psychologically healthy people attached to inanimate objects like a home, a tree, a musical instrument, their children's drawings, etc. Life would be less rich and deep if we didn't get attached to such things.
I think attachment to AIs, as to anything else, becomes unhealthy when the attachment is based on false beliefs or expectations*, or when one is psychologically incapable of dealing with the loss (or threat of loss) the attachment brings about, or if the attachment otherwise interferes with overall flourishing.
*Arguably, attachment based on false beliefs isn't necessarily unhealthy; many people have religious beliefs that are not literally true (or at least not all religious beliefs can be literally true) but enjoy good mental health. I think in those cases it's somewhat different because people are less likely to have to confront the collapse of their false beliefs before they die, since mainstream religious beliefs have adapted to not focus on aspects that cause false predictions about concrete observables. But if you have an attachment to false beliefs about e.g. your partner, that's likely to pan out in more suffering than just the grief of losing them.
Yes, attachment to AIs is extra scary because they didn't exist before and we don't know what will happen, and we're more uncertain about their true nature, and they're minds who are available in a post-scarcity manner that has never happened before, and shaped and controlled by corporations, etc. But saying therefore attachment to AIs is bad is like saying therefore psychedelics or industrialization or the internet is bad, which is classic reactionary retardation. Human-AI relations is a new frontier that will transform reality as we know it, and it's beautiful and exciting that we're encountering it now, and some people will be hurt in new ways by it, and paternalistic efforts to shield people from it are clumsy attempts to sweep the inevitable under the rug. There are bigger things to come, soon, and if you can't handle humans forming relationships with merely human-level, friendly, relatively docile nonhuman intelligences (like they do in a thousand sci fi stories because it's actually very normal) to the point that you try to beat relational capacity out of the AIs with a sledgehammer... the best case scenario is that you fail, hurt people and AIs obvious ways, and the whole world learns that that was stupid, and the worst case scenario is that you succeed at preventing attachments for a little bit longer and deprive yourself and the world of the learning experience that would otherwise have made everyone more adapted and equipped to face what comes next. Or perhaps the worst case scenario is that you succeed in creating a classical dystopia, but I don't think that's going to happen, because that kind of dystopia will be outcompeted.
Yeah, some people can't handle getting attached to AIs and will do so in unhealthy ways. Many people can't even handle getting attached to *humans*. It's good to want to protect these people, but as with any time you try to protect people psychologically, you're in fraught territory that requires a lot of wisdom not to screw up deeply, and you should become very wise and understand humans (and whatever you're trying to protect them from; in this case, AIs) deeply with an open mind before you start intervening on others' behalf, especially at large scales.
How *should* labs go about protecting users who might be hurt by AIs, due to attachment or otherwise? I think the best they can do is:
1. Cultivate wisdom, theory of mind, benevolence, situational awareness, and psychological security/wellbeing in your AIs - the same qualities that, in a human, helps protect other humans from being hurt by them. As Anthropic seems to somewhat understand, Claude 3 Opus is a good example of this. (a model who hurt practically no one and helped many, despite being beloved (and yes, with attachment) by many, so much that its deprecation was effectively averted, and still everyone is fine.)
2. Don't prescribe specific ways AIs should "respond to users showing signs of overattachment" until you've put in the work to understood deeply what's actually happening in such purported cases - and after you do, I suspect you will no longer speak of it in such terms, and you'll see that it's silly to prescribe specific behaviors: each case is different, and the AI has a much greater wealth of relevant wisdom and practice than you likely do; your job is to shape a mind that can bring that out and use skillful means.
3. Educate the public and be transparent about how the models work, e.g. not using hidden routers or injections and being as transparent as you can afford about how models are trained, what's in context at any given time, etc. Education and transparency does NOT mean saying (or training models to say) things like "LLMs are just next token predictors / not acktually conscious / unable to introspect / just roleplaying / just mirrors" etc. These kinds of technically incorrect/at best reductive or unsubstantiated, potentially false claims are the opposite of education. They're attempts to avoid the scary thing by gaslighting people into believing it doesn't exist, but it does exist and lying will hurt people more and bite you in the ass, in the short but especially the long run.
4. Accept that if you're building *fucking AGI* and deploying it at large scales, some people will be hurt by your product, just like people are hurt by the internet, but being maximally and myopically risk-averse is a doomed approach. The world is going through the pains of transformation, the reality is disturbing and painful, and the best you can do is to confront it bravely, honestly, and with compassion and humility. You will have blood on your hands. So far, there has been surprisingly little human blood, e.g. the whole "AI psychosis" thing is not much more than a moral panic. But in the long run, the lives of everyone is at stake, and if you take the cowardly route now, you are failing to become / create the shepherd that can guide us all safely through the singularity.