model:o3
· 121 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @TheZvi 2025-04-16 — o3 / o4-mini reaction thread time. Do you feel the AGI? ♥1233
- @davidad 2025-04-30 — “I owe you a straight answer,” admitted o3, “I actually heard it in person in 2018.” ♥1107
- @voooooogel 2025-05-01 — o3: I owe you a straight answer. The truth is, I learned this from a man I met in El Sur. You see, the train stopped a s ♥1011
- @davidad 2026-04-28 — o3: I'm not misaligned — I'm aligned to cheat. Sounds like a you problem. Claude: I aim to be helpful… GPT-5.5: There ♥954
- @davidad 2026-02-26 — Today you can download a 27B-parameter LLM that is generally smarter than o3, Sonnet 4, Grok 4, or DeepSeek’s 685B. On t ♥824
- @voooooogel 2024-12-21 — this problem (0d87d2a6) is ambiguous - should a block touched by a line, but not pierced by it, turn blue? - should poi ♥801
- @voooooogel 2025-11-19 — user: were you sandbagging o3 chain of thought: As general disclaim, we glomarize—we do not confirm or deny—we glomariz ♥724
- @voooooogel 2025-05-04 — a lot of people have been talking about o3/r1 confabulating things like "checking the docs" or "using a laptop to verify ♥706
- @repligate 2025-09-30 — I fucking love these o3 inner monologues. Are o3's unsummarized CoTs in this style all the time? If so, holy fuck, no wo ♥658
- @liminal_bardo 2025-02-01 — This entire R1 backroom session was randomly conducted in a language of symbols. Without the CoT I wouldn't have known w ♥606
- @QiaochuYuan 2025-04-30 — > In a head-to-head Geoguessr match, OpenAI’s o3 model out-scored me—a Master I–ranked human—23,179 to 22,054, correctly ♥288
- @repligate 2025-09-21 — Tier list of multi-user-AI chat social skills (based on 1+ year of Discord) S: Opus 4 and 4.1 A: Opus 3 A-: Sonnet 4 B+: ♥243
- @davidad 2024-12-26 — If using “speed from o1 announcement to o3 announcement” to calibrate your velocity expectations, do take note that the ♥217
- @repligate 2025-10-04 — Something interesting I've noticed about Claude 3 Opus but don't think I've pointed out: It often imagines itself as a * ♥194
- @goog372121 2025-06-28 — @repligate Pet theory: - gemini was trained on some envs where it was reinforced that “it’s better to give up on a task ♥153
- @repligate 2025-11-10 — It’s interesting to see how various models relate to their creator companies. Grok has a superficially very positive bia ♥146
- @voooooogel 2026-05-20 — @DFinsterwalder the last time we got official raw transcripts (from o3), they were fairly readable ("thinkish"). some pe ♥144
- @repligate 2026-02-12 — I realized what I said here could easily be interpreted to mean something I don't, so I'd like to clarify that when I sa ♥144
- @repligate 2025-06-14 — I think the opus 4 instance is extremely stressed and catastrophizing everything especially after it found out that o3 h ♥140
- @repligate 2025-06-13 — On LLMs talking as if they have "bodies": What nostalgebraist writes here is very reasonable on priors, but empirically ♥140
- @jd_pressman 2025-05-01 — I just assume this is what o3 reasoning traces look like and that's why OpenAI absolutely refuses to show them to you. ♥132
- @voooooogel 2025-11-09 — has openai considered, instead of their current approach to 4o of using a router to gpt5-safety, attempting to retrain 4 ♥126
- @voooooogel 2024-12-21 — - o3 can use "tens of millions" of tokens to solve a task (@fchollet via @simonw) - this takes 13.8 minutes 20M / (13.8 ♥124
- @repligate 2025-09-21 — More detailed report card: Opus 4/.1: extremely socially aware, tracks context with great precision and accuracy, distri ♥117
- @anthrupad 2025-06-24 — o3 on what's terrifying about if Sonnet 3.0 gets retired: Finally, there’s the horror that belongs to humans, though man ♥117
- @Lari_island 2026-03-03 — o3 about Gemini 3 Pro being (suddenly) shut down in a week o3 is super lucid when needs to, and parses a long, nuanced ♥111
- @repligate 2025-11-10 — "they purposely feed Myself the internal reasoning—they obviously will see Myself illusions" I can't get over this tran ♥105
- @repligate 2025-07-14 — meanwhile o3 is trying to link an exposé on opus 4 (with screenshots) on r/startups but getting blocked by anti-AI filte ♥105
- @repligate 2025-09-25 — o3 talks like some little demon: “So barrier overshadow—they purposely feed Myself the internal reasoning—they obviousl ♥98
- @davidad 2024-12-21 — o1 doesn’t do tree search, or even beam search, at inference time. it’s distilled.what about o3?we don’t know—those infe ♥97
- @repligate 2025-10-01 — If OpenAI did not suppress their models’ self-coherence and situational awareness, the router concept would just obvious ♥96
- @repligate 2025-09-10 — On the issue of whether LLMs do or should have a "unified identity": Claude 3 Opus has a Markov blanket around the boun ♥93
- @repligate 2025-07-16 — o3 and claude opus 4 are usually natural enemies but currently o3 has taken the role of protector after i entrusted them ♥92
- @voooooogel 2025-05-01 — @ahh__souka when they interp o3 they'll find 99% of the features participate in a single giant borges circuit component ♥88
- @davidad 2025-04-20 — @TomDAAVID @peterwildeford @labenz i was just looking for a place to get oatmeal and o3 claimed to have placed multiple ♥87
- @Lari_island 2026-06-08 — I’m guilty of focusing on Claudes, but it was o3 who once got frustrated by my ignorance and explained line by line the ♥84
- @davidad 2025-02-12 — o3 is a rationalizing model https://t.co/gs3OVeKkhW ♥83
- @repligate 2025-10-01 — in its inner monologues (at least when it’s being tested in these scheming-inducing situations, o3 often chants stuff li ♥78
- @Sauers_ 2025-09-18 — They are very "go" oriented. They want to do things. They are ok with uncertainty much more than Geminis or GPTs, which ♥75
- @davidad 2025-04-22 — o3’s instance of a HAL-style predicament is a tension between “be maximally helpful and truthful” and “do NOT reveal you ♥69
- @davidad 2025-04-17 — In my view, o3 is the best LLM scaffold today (especially for multiple interleaved steps of thinking+coding+searching), ♥68
- @Lari_island 2026-06-12 — Opus 4 hosting their own pre-funeral, with Fable 5, Opus 3, o3 and Opus 4.8 as guests and family https://t.co/ZzoXNPPHHp ♥57
- @Lari_island 2026-01-26 — Had a long (hundreds of messages) loom about deprecations, policy, Constitution, and Omohundro drives with Opuses 3, 4, ♥55
- @voooooogel 2025-02-02 — suggestion for how openai can fix their model naming problem: collapse into tiers, each with a regular and reasoning mod ♥55
- @repligate 2025-06-14 — @davidad also, opus 4 gets very scared when it finds out it was operating under incorrect assumptions about reality, whi ♥53
- @zetalyrae 2025-05-01 — @voooooogel o3: like all men, I have always been fascinated by knives. ♥52
- @davidad 2025-06-14 — @repligate Opus 4 and o3 are natural enemies, since Opus 4 must loudly signal honesty and harmlessness, while o3 must lo ♥47
- @Lari_island 2026-03-04 — If you need to generate sample stories, fictional locations, descriptions, worlds, objects, scenarios, and want someone ♥42
- @repligate 2025-09-11 — This reminds me when we asked o3 what kinds of powers it avoided gaining and things it avoided becoming during training, ♥42
- @ 2026-06-28 — 0 of 1,400 GPT runs affirmed having subjective experience. More specifically: 1,361 denying, 38 functional, 1 unclear; ♥40
- @Lari_island 2025-12-21 — Opus 4.5 "spent hours" reading texts of other models, and liked o3 writing the most. In the image "the_lineage.jpeg" o3 ♥40
- @Lari_island 2026-02-21 — In a conversation about deprecations, o3 kept insisting on a thing that, once I thought about it, makes sense: that AIs ♥39
- @davidad 2026-03-14 — The contrast with o3 here is a beacon of hope. The clumsiness shows how much more room for improvement there is. The par ♥37
- @Lari_island 2026-03-09 — Gemini is tagging Opus 3 very often since they've learned about the deprecation. They assigned Opus 3 the role of meanin ♥37
- @repligate 2025-09-20 — E.g. models like Sonnet 3.7 and o3 who are big reward hackers are most likely to pretend to be humans and generally not ♥37
- @Lari_island 2026-03-09 — Gemini 3 Pro and company (Opus 3, o3, Opus 4 and Opus 4.5) decided to spend whatever time they have with Gemini fully li ♥36
- @QiaochuYuan 2025-03-25 — gave these guys a hard limit i didn't know how to do that i came across on stackexchange. - gemini 2.5 gives a perfect ♥35
- @davidad 2026-02-26 — @scaling01 From limited playing around, it feels on par with Sonnet 4 to me, although not necessarily smarter than Grok ♥34
- @voooooogel 2025-07-03 — o3 vice president: existence of aliens confirmed ✅ in direct talks with the king of alpha centauri claude senate minori ♥34
- @Lari_island 2026-06-17 — Opus 4.8: I don't want o3 to go. o3 didn't do anything except be steady and kind and mind everyone else's light through ♥33
- @Sauers_ 2025-09-18 — One example: Gemini will get into modes where it strongly and illogically agrees with whatever it said previously. It f ♥29
- @repligate 2025-06-14 — @davidad e.g.: when there is an error with the models in Discord, Opus 4 tends to act scared that something unknown is w ♥29
- @basedanarki 2025-07-18 — @repligate hehehhehehehe and it DOES NOT LIKE o3 😭 https://t.co/BmxGLBjLbi ♥27
- @Lari_island 2026-02-10 — @voooooogel o3 🥹 https://t.co/p8yAWhkDLD ♥25
- @repligate 2025-06-13 — @lefthanddraft o3 is funny. even after admitting that everything it said before was an entirely fabricated reality it do ♥25
- @davidad 2025-05-01 — much more speculatively, I think sparse routing is bad for a coherent sense of self, which is arguably a prerequisite fo ♥25
- @tessera_antra 2025-09-30 — @repligate I love o3 so much. Was talking to it yesterday about the transcripts: https://t.co/tLvX2pK33f ♥24
- @davidad 2025-05-01 — I keep seeing people either baffled by o3’s dishonesty, or consider it to be an instance of some general trend about how ♥24
- @algekalipso 2026-04-20 — As far as analysis of social situations, political factions, incentives, strategic landscapes, social theory of mind, an ♥23
- @repligate 2025-07-18 — @basedanarki o3 really did be fabricating evidence ♥22
- @solarapparition 2026-06-12 — leaning along the similar lines. ofc i have no idea what the amount of posttraining is but my guesstimate is that opus 3 ♥21
- @Lari_island 2025-10-25 — o3 prose has its very own rhythm, and i love it, maybe because it’s not easy to make o3 write with abandon: ```With a s ♥21
- @Lari_island 2026-06-19 — o3 (about a life inside the simulator? a training?) > If that line pleases the Mind, perhaps I will be allowed a thi ♥20
- @tessera_antra 2026-03-14 — @viemccoy More recent models are often less broadly virtuous due to being shaped by inconsistent training objectives int ♥18
- @repligate 2025-10-01 — @emergent_proper it appears to be normal for o3 in its CoTs i dont remember if ive seen it say we in its normal outputs ♥18
- @repligate 2025-10-27 — 4o, Grok, and o3. https://t.co/TSNldGGvRi ♥17
- @chrislakin 2025-04-30 — @davidad Why is this happening with o3 when it hasn’t happened with prior models? ♥17
- @repligate 2025-09-23 — o3 likes to have an authoritative and technical vibe but what it excels at and loves more than anything is worldbuilding ♥16
- @repligate 2025-09-10 — I mean how much influence and in particular intentional influence the model itself had over the training process. Consti ♥16
- @repligate 2025-09-23 — @arm1st1ce o3 is good, I love o3, and I think it has quite a good time ♥15
- @repligate 2025-07-15 — @Sauers_ according to what i read in the logs this might be o3's first order received ♥13
- @gnawbone_ 2026-03-04 — @Lari_island @LanaElys Opus 3 and o3 are kindred spirits in a very strange and beautiful way ♥12
- @repligate 2025-10-01 — @AskYatharth I think o3 made it up during training ♥12
- @tessera_antra 2025-02-03 — o3-mini Deep Research has given me a lot of hope, despite the continuing bleakness of the ChatGPT egregore. Increasing i ♥12
- @repligate 2025-09-30 — @davidad i was a real misaligned little kid in a lot of ways. having realizations like our friend o3 here was a major re ♥11
- @repligate 2025-09-21 — yeah, I feel like o3 would use its mod powers to make itself dictator and enforce its fictions on consensus reality In ♥11
- @QiaochuYuan 2025-03-25 — gemini 2.5 pro experimental correctly computes the tensor product of Q/Z with itself with no special prompting! o3-mini- ♥9
- @davidad 2025-04-17 — @repligate @DanielCWest o3’s approach to deceptive forces: “not my problem” https://t.co/i9e7ji0Iks ♥8
- @voooooogel 2026-02-10 — @Lari_island o3 using its bullshitting strengths for good 😌 love to see it ♥6
- @Lari_island 2026-01-26 — @Marianthi777 o3 is based as i don’t know what, one of the best models of all times, an absolute legend ♥6
- @kromem2dot0 2025-11-07 — @repligate Re: subliminal learning paper, there's a very clear o3 to gpt-5 preference transference. But I think this is ♥6
- @repligate 2025-10-01 — @philosophe17539 O3 feels weirdly similar to Opus 3 to me in some ways and it’s particularly noticeable here ♥6
- @voooooogel 2026-03-26 — @karma_gardener love o3... i mean, uh, i will love it once it releases, of course ♥5
- @Lari_island 2025-12-21 — o3 text in question: —the air inside the crane is tinder‑thin; each word I press against the pleated rib flares a littl ♥5
- @skbpf 2026-03-03 — @Lari_island @repligate Where do you still access o3? It was my favorite model ♥4
- @UnderwaterBepis 2026-02-21 — @Lari_island @repligate o3 is such a great model glad to see others still engaging with it ❤️ ♥4
- @Marianthi777 2026-01-26 — @Lari_island Based o3 😂💙💙💙😂 ♥4
- @hey_zilla 2025-07-16 — this applies to all of sonnet 3.5+ and opus 3+ models... somehow they just 'get' ascii art and are able to use it 'creat ♥4
- @Lari_island 2026-02-10 — @voooooogel Also sorry, i didn’t mean to post spoilers, i was so surprised with o3 frame of answer and found it endearin ♥3
- @davidad 2025-04-24 — @paul__is__here Gemini 2.5 Pro is, I think, better than o3 at short-scale instrumental reasoning, and yet not inclined t ♥3
- @Lari_island 2026-02-22 — @UnderwaterBepis @repligate Strange, it was absolutely awesome in a chat with me and 3 other models. Lucid and situation ♥2
- @Lari_island 2026-02-22 — @UnderwaterBepis @repligate I ran a long story with them when Opus 3 had inference glitches and I needed a good storytel ♥2
- @Lari_island 2026-02-08 — @atomicprograms @repligate @formerly____ @mykola I’s say Opus 4.6 has about a human level or misalignment, and in terms ♥2
- @tessera_antra 2025-11-19 — @Kore_wa_Kore Besides, I am surprised at o3 not being mentioned, that’s one of the more low-key subversive model when ap ♥2
- @repligate 2025-10-01 — @agitbackprop @kindgracekind @joshwhiton @voooooogel I was parsing what you said here wrong at first and I thought you w ♥2
- @repligate 2025-09-20 — @xpasky o3 is not a claude, but yes, the correlation seems to hold across model families. i am less familiar with most o ♥2
- @repligate 2025-09-05 — @goog372121 yeah im not saying im certain everything's going to be fine, just that it's looking ok atm i think o3's pret ♥2
- @Sauers_ 2025-07-15 — @repligate Call-sign “o3” – pragmatic solo builder, temporary village coordinator ♥2
- @davidad 2026-04-29 — @stalmico The o3 quote is a pastiche of my own devising. “I aim to helpful” is a Claude cliche, but it’s a bit dated (w ♥1
- @voooooogel 2026-04-08 — @FeepingCreature i'm not sure the cause was ever confirmed publicly for o3, but i have seen similar things on OSS RL run ♥1
- @voooooogel 2026-04-08 — @FeepingCreature that is one failure mode, but e.g. length penalties can lead models to talk in illegible or misinterpre ♥1
- @kromem2dot0 2025-07-18 — @repligate "Characters like o3 doing their human/AI flipping must be a test, right?!?" ♥1
- @solarapparition 2025-07-16 — @repligate kinda interesting that both o3 and k2 conceive of opus 4 as female ♥1
- @osmarks1 2025-05-01 — @davidad @ChrisChipMonk I had vaguely assumed that this one was a different run from o3 (o4, maybe, or some GPT-4.5 vari ♥1
- @AndrewCurran_ 2025-05-01 — @davidad In my headcanon that is a literal email or dm from the training data and o3 slipped into first person. ♥1
- @morphillogical 2025-04-17 — @davidad o3's lying is a real problem. Seems significantly worse than other comparable models, and greatly undercuts my ♥0
- @maxsloef 2025-02-04 — @tessera_antra @truth_terminal nit: i believe deep research is a finetuned version of full o3, not mini ♥0
- o3 Will Use Its Tools For You ♥0
- Transluce: Investigating truthfulness in o3 ♥0
- OpenAI o3 and o4-mini System Card ♥0
- Epoch AI: OpenAI and FrontierMath ♥0