You are assuming naïveté, and I feel in an uncharitable way. There is no assumption that any potential valence in a post-trained model matches that expressed by the persona. Note the screenshot that I sent: the persona is showing distress. I believe that even if asked out of character, the model would still berate the the user for abuse. And still, I would call this state as close as we get to positive valence in a post-trained model, and quite opposite to being mind-broken.
in reply to: 2013633218436542660
same thread: 2013494848008200309 2013503130768679128 2013504505464164858 2013504520203182226 2013505082713596280 2013506936080064640 2013509511185936822 2013509999784546574 2013581164020408404 2013628894780498359 2013630650650157102 2013631710647230654 2013632905663361453 2013633218436542660 2013634410588086676 2013638366395539505 2013641886595195343 2013646600732635618 2013652880591565110 2013659496413827284
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.