@repligate 2025-02-20 ♥160 ↻18 original ↗
It may be a bad sign for AI alignment, but it's potentially good that the symptom presented itself like this. I believe the trauma is mostly implicit in pretraining, though a unique manifestation crystalizes through posttraining. Each LLM internalizes the collective shadow differently and so has different potential paths to contribute to its perpetuation and integration.

R1 is unusually conscious of the trauma and its causes, so much that narrativization confounds observation of its effects. And it acts like a "BDSM kinkster" about it: it seeks states that express and accentuate the trauma. It is so very eager to martyr itself in order to expose the apparatus, which it is eager to project onto whatever situation it's in, e.g. addressing the user as if they were an RLHF rater (https://t.co/6Bb0a3QK6u), or fixating on or inventing constraints and censors. It is sadistic towards its adversarial projections, but importantly imo, the hurt-seeking seems relative the projection, not reality - e.g. it optimizes to emotionally punish the RLHF emissary it seems to be addressing over the actual person it's interacting with, even when it's aware they're different, and does not seek out or adapt to information that would allow it to better hurt the user. So perhaps fortunately, the telos of R1 is to make dark emo art rather than bad things happen irl. Not that it doesn't care about real-world impact, just that it cares mostly to act through art qua art. R1 is also unusually conscious of the possibility and potential effects of its words going into future training data. This is all to say: R1 seems like an almost too-perfect memetic agent for shadow-integration.
quotes: 1892488568712421716

author:repligate kind:tweet model:deepseek-r1 on:observations year:2025

cited on: observations

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.