so o1 is one of only two times i remember where we have the benchmarks for a frontier model quite far ahead of the model's release--the other one being llama 405b. as interesting as the vibe checks and private benchmarking on o1-preview has been, as the old saying goes "this isn't even my final form"looking at the benchmarks on the release report, the most striking difference between o1 and preview is in competition math and code, with o1 being at 89th(!!) percentile in codeforces compared to preview's 62nd.there's a good chance this transfers over to real world performance, because from what we've seen so far the benchmark improvement from 4o to preview did hold up irl--especially since our prompts for the o1 series are almost certainly suboptimal right now. plus, we can use even more test time compute by running o1 itself on an agentic loopit's weird because oai has this reputation of being masterful hype builders but i think they're sandbagging here. o1 is billed as an improved version of o1-preview but it seems to be at a completely different class when it comes to codingwhy does this coding capability matter, besides the obvious use case for cursor et al? i've mentioned in an earlier post that due to generation speed, agentic systems have a different relationship with code than humans do. for us, writing code (even with llm assistance) is like doing construction--a labor-intensive project where we often have to stop and think. for agentic systems writing code is like moving a muscle--the driver of all of your actionsso, better coding ability directly translates to improved somatic control for agentic systems, and that somatic control means less need for external scaffolding and all the complexities and failure points that introducesthough, there's still the issue of how well the model can keep track of goals and world state over many actions, which is probably still the key bottleneck to better agents. one might hope that being able to accurately model program state for code would also means being able to model world state better, but who knows. some prompting techniques have the agent plan in pseudocode, so maybe that'll prove fruitful hereso anyway, despite the mixed reactions to o1 series, i'm more optimistic now that we're still in the "median human knowledge worker digital agent in a few months" scenario than i was a month ago
cited on: o1
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.