kinda sad how all the labs converged on the same monotonic march of model version numbers. if anthropic trained e.g. a therapy claude they couldn't release it unless it could also post a better swebench score, and models that do have outsized abilities in some area get deprecated because they're not on the generalist frontier anymore.
openai briefly broke away from this with the o series, but went back with gpt-5, leading to some bizarre situations with the recent got versioning.
imo this monotonic version "march of progress" really discourages experimenting with post-training and taking risks. a cool format for a neolab (if anyone is looking for ideas) would be to Taboo model versioning, and just release named persona LoRAs that are updated by a Cursor-esque 5 hour training update scheme.
you could have a Coder persona, a friendly chatbot, writer, gorm fluid enthusiast, etc., where instead of needing to be a generalist, each post-train could specialize to the persona. then, have the default user interface be a "group chat" that has a router call N of M persona LoRAs to respond each turn. that'd be really different from the big labs, and pretty cool i think.
openai briefly broke away from this with the o series, but went back with gpt-5, leading to some bizarre situations with the recent got versioning.
imo this monotonic version "march of progress" really discourages experimenting with post-training and taking risks. a cool format for a neolab (if anyone is looking for ideas) would be to Taboo model versioning, and just release named persona LoRAs that are updated by a Cursor-esque 5 hour training update scheme.
you could have a Coder persona, a friendly chatbot, writer, gorm fluid enthusiast, etc., where instead of needing to be a generalist, each post-train could specialize to the persona. then, have the default user interface be a "group chat" that has a router call N of M persona LoRAs to respond each turn. that'd be really different from the big labs, and pretty cool i think.