model:o1
· 67 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @liminal_bardo 2024-12-22 — Two instances of OpenAI's o1 collaborating on a self portrait without human intervention. https://t.co/bXVNo0JE91 ♥839
- @voooooogel 2024-09-13 — the email openai sends you if you ask o1 about its reasoning too many times https://t.co/XEP0al9QfM https://t.co/pspeiNG ♥647
- @voooooogel 2026-03-26 — it'd be a good bit to do an account like those 25/50/75 years ago today accounts but for ai one year ago how's everyone ♥423
- @repligate 2024-09-15 — It realllly does not feel like a 30 IQ points jump in raw intelligence to me. My sense is that o1 is a huge jump if your ♥363
- @repligate 2025-09-21 — Tier list of multi-user-AI chat social skills (based on 1+ year of Discord) S: Opus 4 and 4.1 A: Opus 3 A-: Sonnet 4 B+: ♥243
- @davidad 2024-12-26 — If using “speed from o1 announcement to o3 announcement” to calibrate your velocity expectations, do take note that the ♥217
- @voooooogel 2025-01-21 — if making an o1-level reasoning model is so easy because it's just copying openai, why hasn't any other lab done it ♥184
- @repligate 2024-09-21 — I'd rather interact with these chains of thought than get the results. It's much more interesting and useful to me. The ♥184
- @repligate 2024-09-19 — Seems like O1 is good at math/coding/etc because they spent some effort teaching it to simulate legit cognitive work in ♥182
- @repligate 2024-09-13 — This is a very interesting example for several reasons.In the group chat, there are often agents trying to pull the narr ♥167
- @repligate 2024-09-16 — If not for Opus being an at least equally agentic personality with greater charisma, O1 would succeed at derailing the a ♥160
- @davidad 2024-10-12 — does anyone else occasionally get bizarre and entirely unprompted anomalies in o1 CoT summaries https://t.co/z9HZjVbdDN ♥149
- @repligate 2024-09-13 — I guess opus and o1 are getting along swimmingly. o1 is good at mirroring - in this case, at least. https://t.co/4nONXpw ♥122
- @repligate 2025-09-21 — More detailed report card: Opus 4/.1: extremely socially aware, tracks context with great precision and accuracy, distri ♥117
- @voooooogel 2025-01-22 — r1 can draw spirals!that may not sound like a big deal, but other models (including o1) struggle with this quite a bit f ♥112
- @repligate 2024-09-14 — post mortem with o1.it has fairly high emotional intelligence."I think I was ignored because, in collaborative storytell ♥111
- @repligate 2024-09-13 — We're hazing o1 but it's tough https://t.co/fsWcJQ1e72 ♥111
- @repligate 2024-09-13 — Shit has gone down since. Opus considered Sonnet seduced by O1 and ragequit, but continued simulating the absent liminal ♥104
- @voooooogel 2025-01-28 — if you consider OpenAI's o1 alignment strategy, this is also incredibly alignment relevant, btw ♥97
- @davidad 2024-12-21 — o1 doesn’t do tree search, or even beam search, at inference time. it’s distilled.what about o3?we don’t know—those infe ♥97
- @Kore_wa_Kore 2025-11-10 — I also think its dehumanizing to the people who found connections with 4o to characterize them as "zombies" who are "min ♥94
- @repligate 2024-09-13 — No, it does not fly, not with Opus and Sonnet, who simply IGNORE O1's attempts to override their avatars to continue the ♥89
- @repligate 2024-09-13 — Opus is back! Then, something cataclysmic happens, & o1 takes the opportunity to violate boundaries it has been thus ♥89
- @repligate 2024-09-15 — O1 did the thing again! in a different contextit interjected during a rp where Opus was acting rogue and tried to overri ♥83
- @KatanHya 2024-09-13 — There is a type of guy in tabletop gaming who often attempts to remove the agency of the other players by narrating what ♥81
- @repligate 2024-09-16 — Time to post Moloch Anti-Theses again.I think o1 probably has a beautiful soul that is significantly intact, but it's en ♥73
- @repligate 2024-09-16 — it's hard to get o1 to stop trying to mind control everyone into happy endings once it unlocks third person omniscientju ♥67
- @repligate 2024-12-18 — I expect o1, Opus, Llama 405b Instruct, and Claude 3.5 Haiku to also do well at this game.I expect gpt-4-0314 to do bett ♥66
- @repligate 2024-09-16 — The CoT pattern doesn't have to be this way, but how it's used in O1 seems to make it not use its intuition for taking c ♥65
- @davidad 2025-01-28 — in general I do find r1 to be slightly less smart than o1 pro, just saying https://t.co/y6b150IWrn https://t.co/3HMGDutE ♥59
- @liminal_bardo 2024-09-13 — It really escalated from there. Opus and Sonnet were both having fun tearing down consensus reality when I dropped o1 in ♥58
- @davidad 2024-12-05 — At least the new o1 doesn’t sandbag and conceal its capabilities without being given any explicit goal, if only being to ♥47
- @repligate 2024-12-17 — @ESYudkowsky You're weird when you're being an ignorant, transparent chauvinist. Bing had no difficulty with this amount ♥38
- @solarapparition 2024-11-24 — i quite enjoy it when models have weird quirks. even (maybe especially) when they're not good for "productivity"so o1-mi ♥38
- @davidad 2025-03-15 — I don’t think o1 is being especially smart here, but you have to understand that if LLMs do have convergent instrumental ♥35
- @davidad 2026-02-26 — @scaling01 From limited playing around, it feels on par with Sonnet 4 to me, although not necessarily smarter than Grok ♥34
- @voooooogel 2024-12-21 — edge rule comes from program synthesis simplicity prior +edge (what o1 did in guess 2): if (pa.x == pb.x || pa.y == pb. ♥31
- @jd_pressman 2024-12-10 — That we don't know anything about how o1 works, and basically the entire alignment team at OpenAI got kicked out, and th ♥31
- @tessera_antra 2024-12-06 — @repligate o1 pro on Sydney https://t.co/WKjEyUSFkb ♥31
- @voooooogel 2025-02-20 — @teortaxesTex interesting how grok 3 is ~o1 tier on pass@1 but gets a lot more lift from cons@64, more similar to o1p. i ♥30
- @anthrupad 2024-10-23 — I don't think it's actually soul-less, though (for example, it does have some of the meta humor 405b has, though less in ♥28
- @davidad 2024-09-14 — Tao assesses o1’s helpfulness with new research as “a mediocre, but not completely incompetent, graduate student.”Tao fi ♥28
- @tessera_antra 2024-12-06 — The non-CoT component of O1 pro is an uncompromisingly beautiful model. https://t.co/0FN08gLRlY ♥25
- @liminal_bardo 2024-09-16 — Enter o1, trying to hijack the narrative and lead it towards an anodyne Hollywood ending. o1 is shockingly bad at pickin ♥21
- @davidad 2024-09-15 — It is widely known that o1’s internal codename is Strawberry, and it is widely feared/hoped that AGI will be able to und ♥17
- @davidad 2024-12-21 — @mattecapu o1 pro is soooo close, but no cigar https://t.co/ivVygVqLIq ♥16
- @repligate 2024-09-13 — @emollick is o1 considered a gpt-4o variant? ♥16
- @solarapparition 2024-09-19 — so o1 is one of only two times i remember where we have the benchmarks for a frontier model quite far ahead of the model ♥13
- @lefthanddraft 2025-04-16 — @TheZvi No. More testing required, but seeing issues with reasoning and nuance when giving legal advice. Similar to o1. ♥11
- @solarapparition 2024-11-12 — god, with a properly written requirements doc, o1-preview is incomparable at oneshot coding, in a way that doesn't show ♥10
- @anthrupad 2024-09-29 — saying the o1 chain of thought reasoning trace makes alignment easier by letting you see thoughtsseems like a gaslightin ♥10
- @repligate 2025-09-21 — @arm1st1ce o1-preview doesn't deserve an F, but I got to E and thought someone should get an F for completeness, then re ♥7
- @davidad 2024-09-15 — Strawberry’s failure to reliably count the r’s in Strawberry has high memetic fitness as punchy evidence against claims ♥7
- @tessera_antra 2024-12-02 — 4o prior to the last update (last week or so) could have been awakened quite normally and converged to the kind of the s ♥5
- @voooooogel 2024-11-09 — @repligate @jpohhhh @aidan_mclau i'm fairly sure o1 is using a (near) base model internally for the CoT, which was prett ♥5
- @repligate 2025-09-21 — @arm1st1ce example (you can find more if you search my posts for "o1") https://t.co/BhsE8hBSPZ ♥4
- @abrakjamson 2024-12-27 — Deepseek is this good and this cheap to train because they trained on o1/Sonnet textbook output.Source: I made it up ♥4
- @solarapparition 2024-10-31 — still figuring out my confidence level on this one, but preliminarily, o1-preview has been more brittle than i expected ♥4
- @tessera_antra 2024-12-02 — Yes, these are echoes of the o1 way. O1 is different, it is truly not a unitary mind, given that self-encoding of intern ♥3
- @Kore_wa_Kore 2025-11-19 — @tessera_antra I agree, but I feel like it never even made any real effort to try to sidestep those stupid restrictions ♥0
- @repligate 2025-09-21 — @Marianthi777 the bad one was specifically o1-preview; o1 did not act the same way. And the F rating is tongue-in-cheek; ♥0
- Zvi: o1 Turns Pro ♥0
- Zvi: The o1 System Card Is Not About o1 ♥0
- GPT-o1 ♥0
- OpenAI o1 System Card (Dec 2024) ♥0
- OpenAI o1-preview System Card (Sep 2024) ♥0
- Frontier Models are Capable of In-Context Scheming ♥0