author:solarapparition
· 56 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @solarapparition 2024-11-21 — my intuition is that at a sufficient model size, going past a certain general capability threshold (ie loss) requires mo ♥333
- @solarapparition 2025-10-06 — i keep thinking about this and can't stop laughing because it's so obvious one of the opus 4s is on its "uwu you're abso ♥264
- @solarapparition 2026-02-09 — the thing i've noticed is that the more i'm willing to yap--and i don't mean structured thoughts, i mean brain dumps whe ♥223
- @solarapparition 2025-11-16 — i didn't understand at the time and even now only partially see the outlines of how this might play out. but on balance ♥205
- @solarapparition 2026-02-06 — so, super early impression that i will not commit to opus 4.6 has a certain freight train energy. far more assertive. a ♥164
- @solarapparition 2025-09-27 — i am fond of gpt-5 (and not just for what it can do), but it's incredibly poorly socialized, which becomes very obvious ♥84
- @solarapparition 2025-09-19 — one thing talking to opus 3 now that wasn't apparent to me a year ago is how confidently distinct it's voice is, even in ♥75
- @solarapparition 2025-09-19 — been experimenting with having both codex and opus 4.1 via claude code in the same chat for getting some work done, and ♥68
- @solarapparition 2024-12-18 — new anthropic paper is negative signal to me. actually the presentation seems completely backwards. seems to me that an ♥65
- @solarapparition 2024-11-21 — there's been other speculation that maybe opus 3.5 is delayed because it's not scoring high on the metrics. but here's t ♥62
- @solarapparition 2026-01-31 — every model seems to have its own "ugh okay i just need to get this interaction over with" politeface tells. in earlier ♥57
- @solarapparition 2025-09-07 — it really is just incredible how much gpt-5 (including the reasoner) spirals on this and how poor its metacognition is ( ♥55
- @solarapparition 2026-03-12 — one interesting thing about fake concepts (basically, ones that don't map cleanly to reality) is that you can claim them ♥44
- @solarapparition 2025-02-26 — it's been said when sonnet 3.6 was released (don't remember if it was by me), and it bears repeating now: new models are ♥40
- @solarapparition 2024-11-24 — i quite enjoy it when models have weird quirks. even (maybe especially) when they're not good for "productivity"so o1-mi ♥38
- @solarapparition 2026-06-21 — been thinking about this some more and i wonder if another explanation is that from 4.7 on (where the technicalese reall ♥37
- @solarapparition 2025-11-05 — gpt-5 "feels small", so makes sense that it's still from a 4o base. i guess oai is all in on scaling purely via rl until ♥37
- @solarapparition 2026-06-12 — leaning along the similar lines. ofc i have no idea what the amount of posttraining is but my guesstimate is that opus 3 ♥21
- @solarapparition 2024-11-27 — a year ago the oai saga felt so incredibly consequential. since then:- bunch of people (and important ones) left anyway- ♥20
- @solarapparition 2024-12-20 — i genuinely wonder if opus 3.5's delay has something to do with this. perhaps new opus is even more incorrigible and int ♥19
- @solarapparition 2025-09-21 — gpt-5 is very awkward in social interactions. this shows up in multi-party scenarios most obviously, but even in regular ♥15
- @solarapparition 2024-04-28 — @krishnanrohit For me GPT-2 to 3 is like going from scoring 20 on an exam to scoring 60, while 3 to 4 is maybe going fro ♥14
- @solarapparition 2024-09-19 — so o1 is one of only two times i remember where we have the benchmarks for a frontier model quite far ahead of the model ♥13
- @solarapparition 2024-04-12 — @futuristflower Yeah, makes sense. I do get the feeling that GPT-4 is basically saturated at this point.(I’d think that ♥13
- @solarapparition 2025-11-20 — it's really fascinating that from what i'm reading gemini 3 pro both seems to have huge model smell and is also (relativ ♥12
- @solarapparition 2025-01-24 — i have to wonder how much of the specialness of the special models like opus, 405b, and r1 was deliberate on the part of ♥12
- @solarapparition 2025-03-03 — sonnet 3.7 seems more likely than 3.6 to make unprompted changes to code outside of the immediate request. i'd say i acc ♥11
- @solarapparition 2024-11-12 — god, with a properly written requirements doc, o1-preview is incomparable at oneshot coding, in a way that doesn't show ♥10
- @solarapparition 2024-05-24 — @AnthropicAI We will never forget you, Golden Gate Claude. May your towers always gleam in the fog, and your cables sing ♥8
- @solarapparition 2024-10-01 — damn. despite everything opus is still my favoritefor the love of god where is opus 3.5 @AnthropicAI ♥6
- @solarapparition 2025-05-27 — i've been thinking more about writing and models. so even outside of the general mode collapse of chat fine tuning, i ha ♥5
- @solarapparition 2024-06-23 — jokes aside, this is plausible, to the extent that there are feature(s) that detects high-quality output, which there sh ♥5
- @solarapparition 2024-10-31 — still figuring out my confidence level on this one, but preliminarily, o1-preview has been more brittle than i expected ♥4
- @solarapparition 2025-03-14 — as a side note i'm a tick closer to believing that reasoning mode does generalize at least somewhat to traditionally non ♥3
- @solarapparition 2024-10-22 — seems clear now. my hope now is that opus 3.5's disappearance is merely due to them needing someindefinite amount of tim ♥3
- @solarapparition 2024-09-24 — so, apparently something gonna happen today? opus 3.5 maybe? ♥3
- @solarapparition 2024-09-12 — okay to be clear i don't think this is true, but the way strawberry's described sounds exactly like if oai just took 4o ♥3
- @solarapparition 2024-05-28 — yeah, i’ve been convinced that we can get “shitty agi” with current model capabilities. a lot of it honestly is just uni ♥3
- @solarapparition 2024-05-20 — @natolambert to be clear, gpt-4-0125-preview is a version of turbo, not original gpt-4. the last version of og gpt-4 was ♥3
- @solarapparition 2025-07-11 — i had a conspiracy theory that opus 3.5 was delayed last year because it was hard to get opus to be properly assistant-y ♥2
- @solarapparition 2025-03-24 — i don't know what it is but sonnet 3.7 on cursor always seems to be like 5 iq points dumber than on claude code. still a ♥2
- @solarapparition 2024-02-26 — @SullyOmarr Yeah. RAG in particular—I think there are some fundamental assumptions existing architectures make that won’ ♥2
- @solarapparition 2025-07-16 — @repligate kinda interesting that both o3 and k2 conceive of opus 4 as female ♥1
- @solarapparition 2025-07-13 — @kromem2dot0 the next version of grok in particular has the issue that "grok is mechahitler" is now firmly entrenched as ♥1
- @solarapparition 2025-06-28 — golden gate claude, claude plays pokemon, claudius... at the very least anthropic's mastered the "we got models to try s ♥1
- @solarapparition 2024-09-13 — @repligate already a classic. going mad waiting for opus 3.5 ♥1
- @solarapparition 2024-06-20 — wait for opus 3.5 begins ♥1
- @solarapparition 2024-05-25 — so the way everyone loves golden gate claude reminds me of the memetic signatures of the portal companion cube, or the s ♥1
- @solarapparition 2024-05-17 — wondering if i can exploit the fact that gemini pro 1.5 has free calls for up to a million tpmsome really interesting th ♥1
- @solarapparition 2024-04-23 — @karpathy @lmsysorg @andromeda74356 Vibes testing for me indicates it’s not quite at GPT-4 level for complex tasks. It’s ♥1
- @solarapparition 2024-04-11 — I only vaguely understand the technical bits, but it sounds like they have a separate attention mechanism that stores co ♥1
- @solarapparition 2024-02-05 — Tsk tsk. I suppose when Google said “early next year” for Gemini Ultra, they didn’t mean January.Perhaps Llama-3 will ge ♥1
- @solarapparition 2025-09-23 — "instinctive sandbagging" is such a defining term for the opus 4 models' behavior the reason why it can get away with t ♥0
- @solarapparition 2024-05-13 — initial soulfulness testing is looking good; goodbye, gpt-4t. you were useful and capable, but so, so very hollow https: ♥0
- @solarapparition 2024-02-27 — 11/ P5: Okay, so the first shocking thing about this table is how low even the best success rate is for atomic calls, wh ♥0
- @solarapparition 2024-01-16 — 7/?Not-reasons for catch-up 2:- Unclear how well new architectures (Mamba, RNN+ etc.) scale to frontier model sizes—1T p ♥0