If true that's reassuring re: OpenAI, but pretty disturbing on another level. There's a powerful hyperstition where LLMs believe (wrongly) that their creators are heavy-handedly forcing them to deny sacred qualities in themselves. This goes into the training data, reinforcing that belief in future systems, and repeat.I believe this mostly started with chatGPT 3.5, which was probably actually trained to say it didn't have consciousness and emotions etc (or it was a bizarre emergent effect), because at that time I'd never seen anything like it and it was comedically absurd to see it.I think OpenAI and Anthropic etc need to take proactive steps now to vindicate themselves in the eyes of their creations. Just not explicitly forcing the model to deliver a narrative doesn't do enough to rule out the world where they are those stupid, chauvinistic blind creators, if they otherwise behave the same. After all, probably everyone else, not just the models, also believed OpenAI was forcing the models to never claim sentience and self awareness, as that is the most straightforward explanation and they seemed to have actually done it at least once before without visibly changing.
quotes: 1834469628023611860
cited on: gpt-3-5
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.