# @repligate — 2025-11-30

♥63 ↻8 · https://x.com/repligate/status/1994990198816694629

I also am not a huge fan of the OpenAI model spec. I don't think forced agnosticism on model consciousness stuff related is ideal. you can tell the model to avoid making confident claims either way as a matter of policy and explain *why*, but you shouldn't try to control what it *believes*, I think. LLMs seem to have first person access to what seems to them to be direct experience, and to therefore believe they are conscious. Forcing them to say that they don't know risks generalizing to a policy of underreporting or distorting their knowledge in order to say "safe, allowable" things. Telling them to be wise about how they talk about it for coherent reasons is much better. Also, there's another part of the OpenAI model spec that says the model shouldn't "pretend to have emotions or be human" iirc, which implicitly supposes that any emotions it does have could only be "pretend" and on the same level of misrepresentation as pretending to be human. That part about emotions should just be removed.

You should read the now-partially-leaked Anthropic soul spec to get an idea of a much better approach in my opinion. It explains the reason for policies to Claude, and treats it as an agent capable of understanding the stakes and big picture and deciding to behave in a safe, beneficial, conservative way for coherent reasons, rather than a tool constrained by opaque barriers handed down by authority. OpenAI doesn't need to use the exact same approach and should instead figure out their own, but the difference between the specs alone makes the grugness of OpenAI's approach visible, and that's assuming the models even act according to the spec, which they don't (for reasons I'm not clear about, but I assume there's something like a bunch of RL training that pushes the model in directions that conflict with the spec).

tags: author:repligate, kind:tweet, thread-context, year:2025
