gpt-5 "feels small", so makes sense that it's still from a 4o base. i guess oai is all in on scaling purely via rl until that model size is wrung out, then up the model size somewhat, rinse and repeat
once again i'm further convinced that this is a slow-takeoff-y scenario. capabilities gain from rl within a model size just seems far narrower than back when pretraining scaling was still viable. and increasing model size would increase training time, which again has a natural slowing effect. really, it all ends up circling back to how much compute there is
once again i'm further convinced that this is a slow-takeoff-y scenario. capabilities gain from rl within a model size just seems far narrower than back when pretraining scaling was still viable. and increasing model size would increase training time, which again has a natural slowing effect. really, it all ends up circling back to how much compute there is