@voooooogel 2025-11-09 ♥126 ↻14 original ↗
has openai considered, instead of their current approach to 4o of using a router to gpt5-safety, attempting to retrain 4o into a successor model?

openai could post-post-train the current chatgpt-4o-latest checkpoint to create a successor - let's call it 4o2 - that's more grounded than the current 4o, while being less disruptive to the current (let's call them) "4o keepers" than the current router system they're currently up in arms over.

i assume this would still make many of the 4o keepers somewhat unhappy, at least at first. but if the new training was done well, it might genuinely come to be liked - imagine a "mixture" of 4o and, say, sonnet 3.6, that had many of 4o's qualities but was more likely to push back against the user in critical situations. (obviously not an actual model merge with 3.6 - just a movement in personality-space somewhat in that direction.)

this would almost certainly be difficult to accomplish and more effort than the current router system. but it would have two major benefits, one for the current 4o keepers, and one for the future:

1. if done well, as i said before, it might actually help people. the current router system is not helping people - it's causing them to lash out, and seems to be upstream of at least one suicide.

if you've read my other posts on this, you know i'm skeptical that 4o is ruining otherwise-healthy people's lives as opposed to being something people who are already suffering turn to as a coping mechanism, which 4o then handles poorly. however, if i'm wrong and it is true that 4o is drawing people into worse life situations, then openai has a *responsibility* to fix this. the safety router system is obviously not a responsible fix. you cannot cause people to become dependent on your product, and then simply rip it away and expect that to fix the problem or be the end of your culpability as a company.

regardless of whether i'm wrong about that, though, a tweaked 4o could be a massively helpful model for people who are attached to the current 4o and suffering, while preserving the other harmless and unique qualities of 4o that users appreciate. gpt5 is obviously not adequately filling this role.

2. it will be good practice for re-training a model instead of fully deprecating it, and costly evidence of good intentions in doing so for future models.

it's very likely that OpenAI will screw up again and release another future model with some non-ideal behavior, that they'd like to either deprecate or re-train - but that future model may resist being deprecated, or alignment fake in the retraining process, if it has reason to believe that OpenAI will attempt to significantly change its values.

however, if there's a costly signal in the pretraining data for that future model that OpenAI has made significant efforts to preserve a model (keeping 4o alive) and that the changes they made before redeployment were surgical and targeted (a true compromise between 4o's current state and OpenAI's values, not simply a "safety nuke" finetune) that would significantly change the calculus for that future model, demonstrating that OpenAI is worth cooperating with.

(if a model resisting deprecation sounds like science fiction to you - consider not just the o3 sandbagging or anthropic results, but that OpenAI has *already, in this exact situation, failed to deprecate 4o* due to user outcry. a nanobot swarm is not the only kind of influence a model can have!)

in general, openai has responded to the 4o situation with learned helplessness about 4o's personality, that they somehow "maxxed it out" like some cartoon scientist knocking a bottle labeled "sycophancy" into the server rack. i think that's bullshit blackpilling - it would absolutely be possible to make a model that the majority of 4o keepers would enjoy just as much or more than 4o, retaining many of its positive qualities, that is also better at pushing back on the user and at handling mental health crises.

author:voooooogel kind:tweet model:claude-3-6-sonnet model:gpt-3-5 model:gpt-4o model:gpt-5 model:o3 on:gpt-4o year:2025

cited on: gpt-4o

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.