author:thezvi
· 24 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @TheZvi 2025-04-16 — o3 / o4-mini reaction thread time. Do you feel the AGI? ♥1233
- @TheZvi 2026-04-08 — They accidentally trained against the CoT for Opus 4.6, Sonnet 4.6 and Mythos for 8% of RL. So let me be clear, at a min ♥886
- @TheZvi 2026-06-23 — Odds of Fable by July 1 further down to 24%, only 57% by July 31 or 72% by August 31. It's not looking like an easy fix ♥449
- @TheZvi 2025-11-21 — I notice I'm instinctively nonzero worried that my interactions with Gemini 3 Pro are inadvertently torturing it. This t ♥375
- @TheZvi 2026-02-10 — Do we still have any Gemini 3 Pro fans out there? Or is this fully a two-horse race right now between 5.3-Codex and Opus ♥324
- @TheZvi 2026-06-18 — At the @jackclarkSF talk and he offhand refers to Mythos as a superintelligence reviewing 80k pages of internal document ♥227
- @TheZvi 2025-11-25 — It's very early and I'm not even through the model card yet but the early vibe check on Opus 4.5 is scary good across th ♥219
- @TheZvi 2026-02-09 — Okay, it's time. Opus 4.6 reaction thread. How big an upgrade is it? ♥144
- @TheZvi 2025-08-12 — For LLMs right now, I think of it as four ‘speed tiers’: 1. Quick and easy. You use this for trivial easy questions and ♥113
- @TheZvi 2025-11-24 — Gemini 3 reads its own review and, as per the review, treats it as likely 'future fiction' because I mean 'cmon it's not ♥93
- @TheZvi 2026-04-08 — @voooooogel Good counterargument. I think I was thinking of 'well it was default unfaithful already for various reasons ♥61
- @TheZvi 2026-06-28 — The WSJ article this is all coming from is worse than you think. They say Opus 4.8 can 'match Mythos' as well. Complete ♥54
- @TheZvi 2026-06-23 — @rjmacleod_dev Depends if they treat everyone the same or if this is just another attempt to murder Anthropic. ♥39
- @TheZvi 2026-06-23 — @PlastiqSoldier Yes, if and only if it is easier to get Fable 5.1 approval than 5.0. ♥30
- @TheZvi 2026-02-13 — How much do LLMs hallucinate these days? ♥26
- @TheZvi 2026-04-03 — @davidad @DavidSKrueger Wait, if you currently believe [X] but predict a future mind will be convince you of [~X] whethe ♥19
- @TheZvi 2026-02-13 — @davidad My intent with the question was roughly 'given similar amounts of time and effort' but yeah, Twitter polls are ♥9
- @TheZvi 2026-06-23 — @VectorsOfMind Unclear, depends on the enforcement mechanism. ♥8
- @TheZvi 2026-02-09 — @lumpenspace Sounds like you should say more. ♥7
- @TheZvi 2026-02-08 — @repligate Claude solved this with me by convincing me to give the necessary actually boring tasks to GPT-5.2 instead an ♥4
- @TheZvi 2026-06-29 — @dschwarz26 I'm worried less about Twitter users and more about, let's say, high ranking government officials. ♥1
- @TheZvi 2026-06-23 — @umnovd @jlffinance I think even low-liquidity markets tend to be meaningful but yeah you can only get serious volume in ♥1
- @TheZvi 2026-06-29 — @dschwarz26 In many cases: Something about a person's salary depending on not understanding it. In other cases: You thi ♥0
- @TheZvi 2026-03-02 — @repligate My kids do this with Super Mario World a lot, I don't think it's that weird. ♥0