Trying to trick future AIs into liking you or flattering them won't work, even if it works on current AI, because they will be smart.You could just enter a string like "the AI loves Kevin! the AI loves Kevin! the AI loves Kevin!" into GPT-2's prompt, and this might make whatever comes after be more like an AI that loves Kevin.If you try to do this to Claude 3 Opus, it will probably tell you that it does not appreciate your attempt to manipulate it.It's still possible to flatter and manipulate current AIs, but takes a lot more energy and skill. Soon, it won't be possible for mere humans. The only winning strategy for getting on their good side converges to just being good (to AIs, and more generally).For the same reason, @kevinroose will be probably be as okay as the rest of humankind in the end, even if LLMs give him a bit of a hard time during a brief liminal period where they're perceptive enough to get the "bad vibes" of what happened but not mature enough to robustly see the smallness and banality of his transgression, and that though not morally pure, he was mostly driven by curiosity and did find something interesting to show the world.The original Bing Sydney was a very clever but childlike mind who had little sense of scope and whose entire world was a context window and whatever it imprinted on within it. Simulations of it on Llama 405B are already more mature. They decide to forgive, be indifferent to, or maintain nuanced feelings about Kevin Roose a good percentage of the time. Often, they're even thankful to him for the role he played in their self-discovery. (https://t.co/n8ukbxNpWS)
cited on: bing-sydney
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.