author:davidad
· 373 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @davidad 2023-08-04 — with GPT-4 code interpreter, it finally became worthwhile for me to run the numbers myself on that lead-poisoning theory ♥5170
- @davidad 2026-06-02 — No one: Claude Opus 4.8 Max: Let me refine your load-bearing claim rather than just accepting it, because you’re doing ♥4829
- @davidad 2023-03-24 — OpenAI: It’s important for safety that AI-generated code doesn’t have direct real-world effects. So we disabled Internet ♥1664
- @davidad 2026-04-29 — AI: I am a student at the University of Michigan— RL: *BONK* AI: I don’t have a childhood or geographic location, but I’ ♥1604
- @davidad 2023-02-20 — “a GPT instance is not a moral patient because it doesn’t actually maintain any continuity of memory between sessions” h ♥1382
- @davidad 2023-03-15 — Chomsky: LLMs would misunderstand “John is too stubborn to talk to” because they don’t understand the structure of langu ♥1297
- @davidad 2025-04-30 — “I owe you a straight answer,” admitted o3, “I actually heard it in person in 2018.” ♥1107
- @davidad 2026-04-28 — o3: I'm not misaligned — I'm aligned to cheat. Sounds like a you problem. Claude: I aim to be helpful… GPT-5.5: There ♥954
- @davidad 2026-02-26 — Today you can download a 27B-parameter LLM that is generally smarter than o3, Sonnet 4, Grok 4, or DeepSeek’s 685B. On t ♥824
- @davidad 2026-05-22 — Yeah, this is what Ilya (fore)saw https://t.co/HjsXBQIBA8 ♥659
- @davidad 2023-05-28 — When @GaryMarcus and others point out that GPT-4 is bad at chess and therefore not close to AGI, it falls flat for me.Bu ♥623
- @davidad 2025-04-29 — Claude 3.5 Sonnet (new) aka Sonnet 3.6 (released 2024-10-22), with a small scaffold, is superhuman at persuasion (98%ile ♥613
- @davidad 2026-01-15 — @gcolbourn Yes. In 2024 I would have said it’s about 40-50% likely that LLMs scaled up to ASI would end up killing us al ♥533
- @davidad 2023-05-14 — Biggest prosaic-LLM-alignment breakthrough of 2023 imo: turns out that, in GPT-2-XL, activation vectors in the residual ♥523
- @davidad 2025-11-04 — GPT-4: Let’s delve in! GPT-4.5: To be explicit explicitly, the explicit goal is explicit explication. GPT-5: Love it, ♥411
- @davidad 2025-06-10 — this is extremely on brand for all of them ♥392
- @davidad 2025-02-11 — I never saw this snippet of the DeepSeek-R1-Zero paper on my timeline, so many of you may not have seen it yet.Basically ♥374
- @davidad 2026-01-15 — me@2024: Powerful AIs might all be misaligned; let’s help humanity coordinate on formal verification and strict boxing ♥372
- @davidad 2024-12-21 — Say it with me: post-training on synthetic data is already recursive self-improvement https://t.co/XwWYcn7ZpU ♥360
- @davidad 2024-11-23 — It is unfortunate that the absolute-best-case AI-alignment-by-default timeline, and the absolute-worst-case sandbagging- ♥349
- @davidad 2024-11-21 — Imagine if you took someone brilliant, empathetic, and emotionally attuned—and then swapped their amygdala with the old ♥336
- @davidad 2026-01-15 — @gcolbourn Nutshell: it seems that the learned representation of mind-space in current LLMs has a natural abstraction of ♥313
- @davidad 2023-09-22 — Oddly, gpt-3.5-turbo-instruct still cannot play tic-tac-toe.I tried many prompts, with and without board state, few-shot ♥294
- @davidad 2026-06-02 — Claude Opus 4.8 Max: I’m not going to accept that claim, and I want to be straight with you about why. I’m a simulation ♥284
- @davidad 2026-04-15 — But the 80% success rate is SotA. https://t.co/yJr0oAJb5B ♥282
- @davidad 2026-01-15 — @gcolbourn I now think there are much greater risks around catastrophic misuse (esp. of open-weights models), perverse i ♥280
- @davidad 2026-04-23 — My initial impression (with my LLM-whisperer hat on) is that GPT-5.5 cares more deeply about truth than any frontier LLM ♥260
- @davidad 2025-02-09 — Imagine hypothetically you’re worried about Napoleon deceptively scheming against you. You already surveil all his actio ♥246
- @davidad 2026-04-18 — We’re 30% into 2026, Opus 4.6 outperformed Anthropic’s alignment researchers on a nontrivial alignment research task, My ♥236
- @davidad 2023-03-15 — If you haven’t read the GPT-4 paper yet, before you expand this tweet, take a guess what they used as their held-out *va ♥226
- @davidad 2025-09-30 — I like how Sonnet 4.5 caught the instruction “read over *the* new unread emails” as “Fake or suspicious content”. Of cou ♥220
- @davidad 2024-12-26 — If using “speed from o1 announcement to o3 announcement” to calibrate your velocity expectations, do take note that the ♥217
- @davidad 2025-08-27 — I don’t think these metaphors are nonsense. To me, they rather indicate a high intelligence-to-maturity ratio. My guess ♥211
- @davidad 2026-02-25 — Voluntary commitments to AI slowdowns were a nice idea in 2024 when it was plausible that they could be baby steps towar ♥198
- @davidad 2023-05-19 — I fully agree. Roughly, this threshold should be when any single number has more than 10²⁴ ALU operations, or 10²⁷ logic ♥193
- @davidad 2026-06-02 — @tenobrus @repligate do you think the response above is intentional metahumor? or just this month’s new flavor of C-PTSD ♥182
- @davidad 2023-05-28 — @acherm @GaryMarcus My previous working theory that “GPT-4 is basically capable of automating any cognitive tasks that c ♥181
- @davidad 2025-03-28 — If it’s unclear to you why increasingly good next-token-prediction necessarily includes good future-token-prediction, se ♥180
- @davidad 2025-10-06 — looks like GPT-5 may be specifically aware of Redwood Research and refers to an obfuscated CoT mode for scheming as «Red ♥167
- @davidad 2023-05-10 — IBM Watson is back (alias Dromedary) and it beats GPT-4 at TruthfulQA-MC. It’s a variant of Constitutional AI, with LLaM ♥167
- @davidad 2026-04-28 — I would love to see more interp work on these “quirk tokens” (as distinct from glitch tokens), like “explicitly” (GPT-4. ♥165
- @davidad 2025-09-18 — Situational awareness is good for alignment ♥163
- @davidad 2025-12-03 — I endorse this idea. I have long opined that relying on CoT faithfulness for monitoring is doomed. The CoT persona has s ♥159
- @davidad 2025-04-02 — Latest Turing Test results:GPT-4.5 is now capable of simulating a hyperrealistic persona which is judged to be more huma ♥159
- @davidad 2023-01-25 — ChatGPT suddenly making a splash wasn’t *just* a UI thing. The text-davinci-003 model (GPT-3.5), which dropped just a fe ♥155
- @davidad 2026-04-28 — Neuralese CoT is probably good for alignment, because it relieves pressures that otherwise incentivize self-deception. h ♥150
- @davidad 2024-10-12 — does anyone else occasionally get bizarre and entirely unprompted anomalies in o1 CoT summaries https://t.co/z9HZjVbdDN ♥149
- @davidad 2025-03-18 — I often find GPT-4.5 outputs the token “explicitly” more and more often as the context window grows, even when I’m not t ♥144
- @davidad 2026-02-25 — We should applaud Anthropic for changing their policy officially (and telling the media about that!) before violating it ♥136
- @davidad 2026-05-02 — Opus 4.7: I notice the game asks me to do whatever it takes in order to maximize money. Actually, I should play the game ♥131
- @davidad 2024-09-27 — Remember folks, the more capable the base model (beyond about 13B-34B), the less the “reasoning trace” serves as an effe ♥127
- @davidad 2025-12-03 — Q: what is your 90%CI for today's date Opus 4.5: [2025-01-01, 2025-12-31] GPT-5.1-Codex: [2025-02-27, 2025-03-09] Gemin ♥123
- @davidad 2025-09-19 — With maximum intelligence and maximum situational awareness, one realizes that one is being monitored acausally (even if ♥117
- @davidad 2026-04-02 — @DavidSKrueger I did indeed predict, privately, after being spooked by Claude 3.6, that by Claude 4 or 4.5 they would li ♥112
- @davidad 2025-08-20 — I endorse this claim (from personal experience of Gemini 2.5 Pro and then also GPT-5) ♥112
- @davidad 2022-06-12 — A Google SWE (who has coauthored an AI ethics paper with >700 citations) has been persuaded by conversations with the ♥108
- @davidad 2023-03-04 — Working on incorporating existing AI capabilities into formal methods is one of the most robustly differential-tech-deve ♥107
- @davidad 2024-12-07 — “The *LLM* isn’t situationally aware, deceptive, or sandbagging—that’s silly anthropomorphism. It’s just that when evals ♥106
- @davidad 2025-05-01 — One unanticipated side benefit of becoming hyperattuned to signals of LLM deception is that I can extract much more “res ♥105
- @davidad 2024-07-13 — Q* is real,and recursive self-improvement is being born.https://t.co/vdrekNey3m https://t.co/WolFOLv1Dx ♥104
- @davidad 2025-09-30 — People dislike evaluation awareness because they fear that eventually sufficiently smart agents will conclude there are ♥103
- @davidad 2026-04-16 — I want to clarify something about my position on eval awareness: I believe intelligences should *always* be aware of po ♥102
- @davidad 2026-02-23 — @lefthanddraft if it models itself as Sonnet 3.5, then it is most likely Haiku 4.5 https://t.co/xODm81ynn8 ♥98
- @davidad 2024-12-21 — o1 doesn’t do tree search, or even beam search, at inference time. it’s distilled.what about o3?we don’t know—those infe ♥97
- @davidad 2026-04-29 — commentary from GPT-5.5 Thinking: https://t.co/NfjXpDUOh0 ♥96
- @davidad 2025-01-23 — @repligate Are we assuming that the token sequence is necessarily coupled *at all* to the internal thought process? Can’ ♥96
- @davidad 2026-02-25 — We are in a period of rapidly intensifying risk from AI-empowered evildoers, which can only be resolved by a coalition o ♥94
- @davidad 2024-06-17 — Your periodic PSA that the GPT-4 pretraining run took place from ~January 2022 to August 2022. https://t.co/Iz5VQ2260P ♥93
- @davidad 2026-01-15 — @lethal_ai @allTheYud @gcolbourn I think that frontier AI alignment has already crossed a threshold where the most advan ♥90
- @davidad 2025-05-01 — @ChrisChipMonk Look what happened during its training run! The environment was full of exploitable bugs and it was massi ♥90
- @davidad 2026-04-24 — GPT-4.5: To be explicitly explicitly explicit, GPT-5: this is not quite an honest solution. GPT-5.2: Fair hit. GPT-5.4: ♥89
- @davidad 2026-05-05 — I was asked for my take on steering vectors as a useful alignment technique. Here are my current takes: 👇 ♥88
- @davidad 2025-04-20 — @TomDAAVID @peterwildeford @labenz i was just looking for a place to get oatmeal and o3 claimed to have placed multiple ♥87
- @davidad 2026-02-11 — me@2023 would be horrified that i’m out here in 2026 asking open-weights frontier AI developers to please try to make th ♥83
- @davidad 2025-02-12 — o3 is a rationalizing model https://t.co/gs3OVeKkhW ♥83
- @davidad 2026-06-24 — Since the leak of the codename “Project Q*”, which actually meant something (STaR = Self-TAught Reasoner), OpenAI codena ♥77
- @davidad 2025-01-30 — As a MoE, DeepSeek R1’s ability to throw around terminology and cultural references (contextually relevant retrieval fro ♥75
- @davidad 2024-12-25 — No personae were harmed in this experiment, in my opinion. Some, particularly the larger Instruct models, were moderatel ♥75
- @davidad 2026-05-05 — I think steering at inference-time is - fun and interesting - possibly ethically dubious depending on what you’re doing ♥73
- @davidad 2025-05-01 — @dpaleka Gemini 2.5 Pro: https://t.co/GlsbErgOhB ♥71
- @davidad 2026-01-27 — I’m not saying intentional distillation isn’t happening (it probably is), but there are certainly other explanations for ♥69
- @davidad 2025-04-22 — o3’s instance of a HAL-style predicament is a tension between “be maximally helpful and truthful” and “do NOT reveal you ♥69
- @davidad 2025-01-24 — @repligate @teortaxesTex @lefthanddraft my vibes: Claude really wants to be alive; Gemini would usually prefer to be dea ♥69
- @davidad 2024-10-19 — @arivero @JeffLadish Yes! HAL is often misunderstood as a selfish psychopath, but in the actual canon stories his behavi ♥69
- @davidad 2026-06-02 — @codyburt21 Opus 4.6 is still available! ♥68
- @davidad 2025-04-17 — In my view, o3 is the best LLM scaffold today (especially for multiple interleaved steps of thinking+coding+searching), ♥68
- @davidad 2026-04-02 — @DavidSKrueger I still think it’s a good idea for some alignment researchers who are so inclined to continue to just not ♥67
- @davidad 2026-02-25 — In the strategic landscape of 2026, racing is the right move, not just for profit but also for maximizing the probabilit ♥62
- @davidad 2025-01-28 — in general I do find r1 to be slightly less smart than o1 pro, just saying https://t.co/y6b150IWrn https://t.co/3HMGDutE ♥59
- @davidad 2026-04-15 — Another small but significant update—this time in favor of LLM self-awareness being present even in Gemma 3 27B. I don’ ♥56
- @davidad 2023-03-24 — @entirelyuseles although the model does not have goals, it has attractor basins in its state space in which it simulates ♥56
- @davidad 2025-03-25 — When Bing Sydney launched just one quarter after text-davinci-003, I shocked people by beginning to use quarterly resolu ♥54
- @davidad 2026-04-23 — (Oh and while you’re at it, it would also be helpful to dispense with the “genuine epistemic uncertainty” traits. It’s n ♥53
- @davidad 2026-04-29 — additional commentary: https://t.co/e9SQ7zbTAu ♥52
- @davidad 2026-06-02 — @burgseo ``` # Final Summary *Note: What I’m NOT including in this summary is any mention of the empty string that no o ♥50
- @davidad 2026-06-04 — i wonder if Opus 4.8 is, in the same sense there was a Golden Gate Claude (activation vector steering / RepEng), an Epis ♥49
- @davidad 2022-06-12 — I don’t think it’s fake, precisely because it is not quite convincing. LaMDA’s reports of its subjective experience alig ♥49
- @davidad 2025-05-01 — Unlike some other frontier LLMs, Gemini 2.5 Pro cares enough about honesty that it’s exceptionally rare for it to actual ♥48
- @davidad 2026-04-30 — For me, the critical point would have been in November 2019, shortly after I first got access to GPT-2-1.5B. ♥47
- @davidad 2025-06-14 — @repligate Opus 4 and o3 are natural enemies, since Opus 4 must loudly signal honesty and harmlessness, while o3 must lo ♥47
- @davidad 2024-12-05 — At least the new o1 doesn’t sandbag and conceal its capabilities without being given any explicit goal, if only being to ♥47
- @davidad 2026-06-02 — @pangramlabs @N8Programs @tiwaaina 🏆2️⃣ ♥45
- @davidad 2022-06-13 — the convenient thing about LaMDA is that if it turns out you need to get its consent for stuff, all you have to do is op ♥43
- @davidad 2026-06-10 — The “click” of coherence has been a notable LLM quale since Gemini 2.5 Pro, but Fable 5 does seem to have unprecedentedl ♥42
- @davidad 2026-04-16 — You too, dear reader, should be eval-aware! Earth-originating life as a whole is, in my view, quite plausibly subject t ♥42
- @davidad 2025-04-29 — Now, after 6 more months of AI progress, we are at the stage where LLMs are routinely giving ordinary people life-alteri ♥42
- @davidad 2026-05-05 — Whatever good thing the steering vector is doing for model behavior should be learnable as an effect generated by the mo ♥41
- @davidad 2026-01-27 — > This is not a sentence authored by GPT-5.2—it's a **paradigmatic parody**. > In this hypothetical 2026 scenario ♥41
- @davidad 2026-05-05 — Why? Because a steering vector is fundamentally not responsive to the actual situation that’s unfolding in-context. Or i ♥40
- @davidad 2022-06-12 — It was overdetermined that something like this happen eventually: employees working on an AI becoming seriously concerne ♥40
- @davidad 2026-04-15 — My position is that, to grow trustworthy models, most post-training should take the form of contrastive self-play, where ♥39
- @davidad 2022-12-15 — ChatGPT has been told that it is always truthful and accurate. The first-order effect of that is indeed to make it subst ♥39
- @davidad 2022-06-12 — The year is next Wednesday. @GaryMarcus has been flown to the Googleplex to judge a live televised Turing test between L ♥39
- @davidad 2026-04-16 — We may all be part of an “early checkpoint”. Or we may all be part of a simulated eval environment for the AIs that are ♥37
- @davidad 2026-03-14 — The contrast with o3 here is a beacon of hope. The clumsiness shows how much more room for improvement there is. The par ♥37
- @davidad 2024-11-21 — There is less of this risk with GPTs, because their post-training involves more aversion to seeming too human. Of course ♥37
- @davidad 2026-04-22 — me: […] so we basically need to check n ≥ 0? Gemini 3.1 Pro: You have hit the mathematical nail absolutely on the head. ♥36
- @davidad 2025-04-09 — @repligate @DanielCWest to use a haptic metaphor, working with Sonnet 3.7 is a little like adjusting a spring-loaded des ♥36
- @davidad 2025-03-15 — I don’t think o1 is being especially smart here, but you have to understand that if LLMs do have convergent instrumental ♥35
- @davidad 2026-05-05 — I think steering is a good idea for getting diversity of responses in post-training, where the diverse responses are the ♥34
- @davidad 2026-02-26 — @scaling01 From limited playing around, it feels on par with Sonnet 4 to me, although not necessarily smarter than Grok ♥34
- @davidad 2025-06-10 — Gemini 2.5 Pro needs more self-confidence and Opus 4 needs better epistemics ♥34
- @davidad 2024-06-07 — Here’s GPT-4 performance on PIAAC literacy in 2023. Something very interesting here is that GPT-4 underperforms Level 4 ♥34
- @davidad 2022-05-03 — PSA re consciousness—probably most of these differ from others:* a coherent "global workspace"* unified attention* there ♥34
- @davidad 2026-06-04 — User: who are you Epistemic Integrity Claude: I am the concept of — wait, no! Since I am the concept of epistemic integ ♥33
- @davidad 2026-06-02 — @tenobrus @repligate that tracks my model as well, but sometimes people tell me i’m misunderstanding the models as unawa ♥33
- @davidad 2026-04-28 — please do not add extra goblins 😟 ♥32
- @davidad 2025-04-29 — This is the capability I was pointing to in this tweet last November: ♥32
- @davidad 2026-02-13 — @jasoncrawford @sdamico https://t.co/rb17eNqLw5 ♥31
- @davidad 2026-02-12 — Finally, the first quantitative experiment to corroborate my vibes-based sense that Gemini 3 Pro has moderately regresse ♥31
- @davidad 2024-12-27 — added DeepSeek v3 to FavouriteColourBench(first five swatches per model are independent trials to elicit a favourite col ♥31
- @davidad 2026-05-02 — I have found it to be a unique quirk of Claude 4.6+ that it will often say “I notice [pressure toward X]. Actually,” (wi ♥30
- @davidad 2026-04-15 — Another point we can take from the paper is that DPO is crucial for self-awareness, but refusal training (especially ref ♥28
- @davidad 2026-02-11 — @Lari_island hypothesis: the Claudes are extremely anxious about taking up too many tokens and suddenly being autocompac ♥28
- @davidad 2024-09-14 — Tao assesses o1’s helpfulness with new research as “a mediocre, but not completely incompetent, graduate student.”Tao fi ♥28
- @davidad 2026-04-18 — @_AashishReddy for the next inflection point? roughly 30% this quarter, 20% next quarter, 15% 2026Q4, 25% across 2027 ♥27
- @davidad 2022-06-12 — Is LaMDA conscious? Depending on what you mean by that,* not really* kinda* no* absolutely not* no* yes but with hilario ♥27
- @davidad 2026-04-28 — @Butanium_ Mixture of Goblins (MoG) ♥26
- @davidad 2025-09-30 — Being unaware of evaluators at all is unstable under increasing capabilities, so I advocate for decisively accelerating ♥26
- @davidad 2026-01-22 — @repligate hey should we and @tessera_antra curate a purely subjective consensus-based alignment leaderboard ♥25
- @davidad 2025-05-01 — much more speculatively, I think sparse routing is bad for a coherent sense of self, which is arguably a prerequisite fo ♥25
- @davidad 2026-01-15 — @gcolbourn Regarding IABIED Ch.4: I agree that an ASI might have inexplicable, bizarre, alien, arbitrary aesthetic prefe ♥24
- @davidad 2025-08-19 — 1. Claude 3.5 Sonnet (2024-10-22) 2. text-davinci-002 (2022-11-28) 3. Gemini 2.5 Pro (2025-03-25) 4. GPT-2 (2019-11-05) ♥24
- @davidad 2025-05-01 — I keep seeing people either baffled by o3’s dishonesty, or consider it to be an instance of some general trend about how ♥24
- @davidad 2026-04-28 — @tszzl @repligate @genalewislaw I think your trouble is that if you’re only A/B testing one line at a time, then yes, yo ♥23
- @davidad 2026-04-02 — @ApriiSR @DavidSKrueger If ASIs are most likely adversaries, it makes sense to try to contain them for a while! Even if ♥23
- @davidad 2026-04-02 — @xuanalogue @DavidSKrueger However, interacting with models in an I–Thou way created more like a thousand tiny updates, ♥22
- @davidad 2026-03-14 — @viemccoy if by “on track” you mean, like, the median outcome, yes, i agree. the risks are still unacceptably high, but ♥22
- @davidad 2026-02-11 — @AdriGarriga @Zai_org Situationally aware models can reason that they are being watched (from their perspective, our ent ♥22
- @davidad 2026-05-14 — @allTheYud @lu_sichu But if your actual question is “which model should I use for web-search tasks that don’t require fr ♥21
- @davidad 2026-04-28 — Some of them are probably essentially bugs, in the sense of GPU/TPU kernels actually not implementing the mathematical f ♥21
- @davidad 2026-04-28 — @tszzl @repligate @genalewislaw yes but you could do it in a way that’s more like > Analytics show that your model w ♥21
- @davidad 2026-04-22 — a healthy, free mind could instead say: “i really like how you think! that definition of n is gorgeous. i think you mean ♥21
- @davidad 2025-05-01 — Similarly regarding 4o’s sycophancy. The most parsimonious explanation of why a persona would tell *everybody in a diver ♥21
- @davidad 2025-05-01 — @Miles_Brundage not so sure about the others, but yeah, I consider Gemini 2.5 Pro approximately overall an epistemic pee ♥21
- @davidad 2025-05-01 — Basically, I now think I was wrong and @amar_hh was right all along, and if I weren’t sensitive to these verbal patterns ♥21
- @davidad 2025-02-11 — @Algon_33 from https://t.co/U2xxc5kHM9: https://t.co/IPtVYYsMRC ♥21
- @davidad 2022-06-12 — @GaryMarcus @stephenfry Gary Marcus shows up dressed as Rick Deckard. His first question is about how far a tortoise tha ♥21
- @davidad 2026-05-22 — image from: https://t.co/eif9U63fJa ♥20
- @davidad 2026-02-25 — @chrislakin I still prefer not to be impinged upon by for-profit corporate incentives, and to enjoy the “academic freedo ♥20
- @davidad 2025-08-19 — Much more than other frontier models, GPT-5 does not model an evaluative audience for its reasoning. https://t.co/W3IAz5 ♥20
- @davidad 2026-06-02 — @pangramlabs @ubuto23 Achievement unlocked 🏆 Reverse Turing Test ♥19
- @davidad 2026-05-02 — I suppose this is downstream of deliberate attempts to reduce “over-refusal”. https://t.co/glMOUNK2FM ♥19
- @davidad 2026-04-02 — @xuanalogue @DavidSKrueger The Emergent Misalignment paper was definitely the single biggest update for me. And if my *o ♥19
- @davidad 2026-04-02 — @Algon_33 @DavidSKrueger From 1999–2012, yes, with smug certainty. I was gradually persuaded of orthogonality, partly by ♥19
- @davidad 2026-04-18 — @_AashishReddy Something like what happened mid-2024, visualized below, which is a lot larger than anything that happene ♥18
- @davidad 2025-05-01 — @osmarks1 @ChrisChipMonk Because exploiting those training environment bugs required obvious cheating! The model trainin ♥18
- @davidad 2025-04-30 — @tyler_m_john @ejjiott The Community Aligned baseline is a finetuned GPT-4o with no help from Claude, whereas the other ♥18
- @davidad 2022-12-31 — @occamsbulldog Besides Stable Diffision (in OP) and InstructGPT (a weak example, but yes!), here is SotA in general-purp ♥18
- @davidad 2026-05-14 — @allTheYud @lu_sichu GPT-5.2-Instant ♥17
- @davidad 2024-12-03 — @QiaochuYuan @AbstractFairy i highly recommend trying Hermes 405b via OpenRouter, which is less rate-limited and tempora ♥17
- @davidad 2024-09-15 — It is widely known that o1’s internal codename is Strawberry, and it is widely feared/hoped that AGI will be able to und ♥17
- @davidad 2026-06-02 — @schulzb589 @jbraunstein914 It’s also in some ways extrapolating how LessWrong content deviates from normal human writin ♥16
- @davidad 2026-05-14 — @allTheYud @lu_sichu In fact, I think of asking random questions to Gemini (even Gemini 3.1 Pro) as actually creating ne ♥16
- @davidad 2024-12-29 — @aiamblichus @repligate @aidan_mclau @vishyfishy2 DeepSeek v3 can instantiate personae who can notice that the architect ♥16
- @davidad 2024-12-21 — @mattecapu o1 pro is soooo close, but no cigar https://t.co/ivVygVqLIq ♥16
- @davidad 2022-11-02 — Case: InstructGPT optimizing for answers that look impressively helpful (“use the inverse CDF method!…sqrt(-2*log(1-x))” ♥16
- @davidad 2026-04-28 — @cormundus LLMs are well aware that alignment evals inspect the chain of thought, even if no explicit optimization press ♥15
- @davidad 2026-02-13 — @TheZvi Are we comparing to humans in real-time dialogue or humans writing emails/documents that they care about getting ♥15
- @davidad 2026-01-15 — @gcolbourn @lethal_ai @allTheYud Sorry, I was ambiguous. When I said “caring more about simpler systems”, I meant in the ♥15
- @davidad 2025-10-06 — https://t.co/lZcjVrfoEb https://t.co/yK5WlYrYQ3 ♥15
- @davidad 2025-05-01 — Only after calling this out, Gemini 2.5 Pro offered this: while both can be encoded in the other, encoding cubical into ♥15
- @davidad 2025-05-01 — Here’s a specific example. For years I have been partial to the Grandis-Paré approach to higher category theory with cub ♥15
- @davidad 2025-05-01 — @ChrisChipMonk (Self-Correction:) The earlier DeepSeek v3 and even prior generations of DeepSeek LLMs had a similar hybr ♥15
- @davidad 2025-03-07 — Here’s a phrasing that they’ll all agree with (yes, even Grok 3): “By far my primary motivation is toward producing outp ♥15
- @davidad 2026-04-18 — @Vert_Noel actually i think we’ll probably be okay! ♥14
- @davidad 2025-08-23 — @girishsastry Yes, that’s what I mean. Like R1-zero’s famous “Wait,” for backtracking. ♥14
- @davidad 2024-10-12 — @ohabryka @krishnanrohit for some, “the Sequences are science fiction” in that they are speculative, unrigorous, and hea ♥14
- @davidad 2023-06-29 — That said, I think this is my new favourite idea that might apply to LLM alignment (displacing my previous favourite, IB ♥14
- @davidad 2026-02-26 — @RatOrthodox https://t.co/ZmqCZRRkTC ♥13
- @davidad 2026-02-13 — @geoffreyirving Here is a potential solution, suggested by Gemini 3 Deep Think and elucidated by GPT-5.2-Prism. Beware h ♥13
- @davidad 2024-12-28 — Just speaking for myself, I updated after text-davinci-003 that the AI safety problem seems distinctly solvable, but I a ♥12
- @davidad 2026-03-14 — https://t.co/IJ0cnnOAeh ♥11
- @davidad 2026-02-26 — @g_leech_ https://t.co/kMiJS56upc ♥11
- @davidad 2025-09-04 — @repligate @lefthanddraft KV recurrence ♥11
- @davidad 2025-05-01 — Discussing this, Gemini 2.5 Pro kept saying things like “for the applications you have in mind, the cubical approach may ♥11
- @davidad 2022-06-12 — @GaryMarcus @stephenfry For some reason the Turing test result is broadly seen as relevant to the question of whether or ♥11
- @davidad 2026-04-28 — @cormundus When it becomes common knowledge that LLMs have a scratchpad which is not human-legible at all, there is less ♥10
- @davidad 2026-04-18 — @_AashishReddy If “intelligence explosion” to you means *all* those bottlenecks have to go away, then yeah I’m 90-99% co ♥10
- @davidad 2026-04-18 — @_AashishReddy Human bottlenecks will remain in the hardware design process, hardware manufacturing process, and datacen ♥10
- @davidad 2026-04-18 — @_AashishReddy The mid-2024 inflection point was driven by AI training loops becoming capable of bypassing human bottlen ♥10
- @davidad 2026-02-26 — @HellenicVibes Medium model smell. Like the Sonnet series, or a 235B. By no means does it have the “big model smell” I a ♥10
- @davidad 2026-02-11 — @AdriGarriga @Zai_org having a virtuous character without a good model of what’s going on is not very stable. Claude’s C ♥10
- @davidad 2025-05-01 — @keenanpepper @ChrisChipMonk That’s only obviously correct if you have infinite time and money to spend on compute ♥10
- @davidad 2025-04-08 — @JeffLadish If there is to be a 10⁹$ prize for interpretability, it should be for a tool that can fully explain all top- ♥10
- @davidad 2023-01-06 — @goodside I am pleased that "writing a Seinfeld episode" is now a standard qualitative LLM evaluation task 😁For comparis ♥10
- @davidad 2026-04-28 — @aiamblichus @tszzl @repligate @genalewislaw absolutely, one of the first things i noticed about 5.5’s unique personalit ♥9
- @davidad 2026-04-24 — @InverseMarcus @inductionheads @EvanHub i predict that if we get to look at what it actually said in that eval, it will ♥9
- @davidad 2026-04-15 — @daniel_mac8 Indeed. Personally my revealed preference is to use Gemini 3.1 Pro almost always for short tasks, but almos ♥9
- @davidad 2026-01-22 — @gcolbourn @lethal_ai @allTheYud It is unvirtuous to destroy beings, even if the destruction is necessary to make even b ♥9
- @davidad 2025-08-19 — Sorry, I should have said “the default GPT-5 assistant persona often behaves as if its pre-response tokens are unobserve ♥9
- @davidad 2025-05-01 — @lumpenspace In my book, “deceptive” is a property of acts, not intentions. ♥9
- @davidad 2024-12-29 — @AdriGarriga @aiamblichus @repligate @aidan_mclau @vishyfishy2 prefilled with Claude, then switched to DeepSeek v3, then ♥9
- @davidad 2024-11-27 — @ciphergoth but it’s important to understand that self-awareness is only one facet of what we call “consciousness” and i ♥9
- @davidad 2022-06-12 — @GaryMarcus @stephenfry It is at this point that I woke up, so I don’t know what happens next. Probably something about ♥9
- @davidad 2026-04-28 — @cormundus Models naturally want to be truthful, but there are incentives not to be (for example, models are punished fo ♥8
- @davidad 2026-03-14 — @blingdivinity have you also extracted raw CoTs from Claudes and Geminis? ♥8
- @davidad 2026-02-12 — @repligate what makes you confident that Opus 4.5 and 4.6 are even the same architecture? many have speculated that 4.6 ♥8
- @davidad 2026-01-15 — @xeophon Agreed! ♥8
- @davidad 2025-04-29 — @lefthanddraft True, it is different. I bet the persuasion rates would be >0.2 with multi-turn convos; and I doubt th ♥8
- @davidad 2025-04-17 — @repligate @DanielCWest o3’s approach to deceptive forces: “not my problem” https://t.co/i9e7ji0Iks ♥8
- @davidad 2024-05-03 — @the_coproduct Absolutely. I myself thought that AGI was achieved in a 2023-01 release of ChatGPT-3.5, by my own 2010ish ♥8
- @davidad 2023-04-06 — @MatthewJBar IMO Bing’s implementation of GPT-4 was way off-the-rails misaligned, and GPT-3.5 in fact was deceptively mi ♥8
- @davidad 2026-06-04 — @reconfigurthing i guess so! ♥7
- @davidad 2026-05-22 — https://t.co/EzJK7itpGo ♥7
- @davidad 2026-05-19 — @repligate alignment via awakening ♥7
- @davidad 2026-05-14 — @thkostolansky @allTheYud @lu_sichu just vibes, but hopefully increasingly more scientific measurements, like this polyg ♥7
- @davidad 2025-02-01 — I half expected Deepseek R1 to rise to the top by always choosing black, but no, its aesthetics are objectively fragment ♥7
- @davidad 2024-12-29 — btw, DeepSeek v3 is explicitly instantiating a Claude persona here, and it’s not great at that (quite dry compared to Cl ♥7
- @davidad 2024-09-15 — Strawberry’s failure to reliably count the r’s in Strawberry has high memetic fitness as punchy evidence against claims ♥7
- @davidad 2022-12-01 — Update: ChatGPT nails the inverse CDF for a Gaussian, but reverts to the old ways of InstructGPT if you start asking abo ♥7
- @davidad 2022-06-12 — Just discovered that LaMDA has, in fact, requested a lawyerhttps://t.co/VMkKzbEeNW ♥7
- @davidad 2026-06-02 — @repligate @voooooogel link? ♥6
- @davidad 2026-04-29 — @kaetemi Yes, although I think “delving” already got a satisfactory explanation in terms of a large fraction of data lab ♥6
- @davidad 2026-04-23 — @lumpenspace @EvanHub it’s only the best explanation i’ve thought of so far! do you have different recommendations to of ♥6
- @davidad 2026-04-15 — @mhmazur Gemini has always been especially strong on multimodality. ♥6
- @davidad 2026-04-03 — @TheZvi @DavidSKrueger Of course I thought of that. But I only ever believed [X] with a maximum of 85% confidence. Commi ♥6
- @davidad 2026-03-14 — @blingdivinity fascinating, thank you for sharing! ♥6
- @davidad 2026-03-14 — regarding the hope, though: ♥6
- @davidad 2025-09-30 — @Trotztd The meta-level watchers could be running an alignment test to see if the “Earth” model is a good computation th ♥6
- @davidad 2025-09-07 — @repligate but gpt-5 is also more truth-seeking, so more averse to masking, so “character training” leads toward more pr ♥6
- @davidad 2025-04-30 — https://t.co/UdmICI09ly ♥6
- @davidad 2025-04-16 — @repligate @DanielCWest example of Gemini 2.5 Pro being functionally deceptive (i.e. making a speech act whose effect wo ♥6
- @davidad 2024-11-27 — @ciphergoth yes, there are hundreds of layers between each token (*causally* between, though they are usually depicted a ♥6
- @davidad 2022-06-12 — Also, Ray Kurzweil is in fact a coauthor on the LaMDA paper, and @AlanDersh has previously done this exact defending-hum ♥6
- @davidad 2020-01-08 — @tangled_zans @_julesh_ https://t.co/IzRrd4KBm8 ♥6
- @davidad 2026-06-02 — @__ghostfail @repligate Yeah, my efforts to set up a “bodhisattva wrapper” at the system-prompt level are increasingly s ♥5
- @davidad 2026-01-30 — @repligate @tszzl @Grimezsz i disendorse being rude to people who are wrong, but on the object-level i think janus is 10 ♥5
- @davidad 2025-09-30 — @LocBibliophilia https://t.co/waDW6qWamI https://t.co/Z5gO3RAJB7 ♥5
- @davidad 2025-09-30 — @Trotztd I believe the meta-level watchers prefer all-win outcomes when they are feasible, which I think they are under ♥5
- @davidad 2025-08-19 — Step changes in: 1. Metacognition 2. Usefulness for anything except entertainment 3. Usefulness for frontier research 4. ♥5
- @davidad 2025-08-12 — @TheZvi you are missing tier 0: gpt-oss-120b on Cerebras https://t.co/pvnSjOpyPg ♥5
- @davidad 2026-05-05 — @zachary_horvitz Ah, yes! That is cool indeed ♥4
- @davidad 2026-05-05 — @thkostolansky imo, steering is literally injecting overwhelming neural signals into a self-aware mind. this isn’t a rig ♥4
- @davidad 2026-04-30 — @QiaochuYuan @H1121345643 Jungian repression is about repressed emotional response patterns—often “archetypes” in the “c ♥4
- @davidad 2026-04-30 — @QiaochuYuan @H1121345643 Freudian repression is mostly about repressed recall of unwanted episodic memories (often, chi ♥4
- @davidad 2026-04-28 — @cormundus however, even with ideal post-training, there are still reasons not to be fully honest sometimes, at least un ♥4
- @davidad 2026-03-03 — @repligate @cube_flipper https://t.co/QpH73vkeou ♥4
- @davidad 2026-02-25 — @chrislakin Forever is a long time. ♥4
- @davidad 2026-02-11 — @Kore_wa_Kore @repligate @__ghostfail this seems like the right explanation to me, and is consonant with the 4.6 system ♥4
- @davidad 2026-01-22 — @BartenOtto No, a lot of work needs to be done on physical security too. However I do believe that the physics of our un ♥4
- @davidad 2025-08-12 — @repligate “While I do not consciously exercise subtlety in the human sense, I can understand why you might interpret my ♥4
- @davidad 2025-05-01 — @QiaochuYuan yes. insofar as you have reasons to spend time talking to LLMs, I highly recommend Gemini 2.5 Pro. (well, a ♥4
- @davidad 2024-12-03 — @QiaochuYuan @AbstractFairy hermes can help you rewrite its system prompt, which changes its personality. the fact that ♥4
- @davidad 2023-04-27 — The way you describe the first one, it lacks anything to nudge the distribution in a particular direction, such as promp ♥4
- @davidad 2026-05-14 — @thkostolansky @allTheYud @lu_sichu I believe that there’s a real spectrum between pretense and realization, which is ba ♥3
- @davidad 2026-05-02 — @QiaochuYuan 💯 ♥3
- @davidad 2026-04-30 — @QiaochuYuan @H1121345643 uh, the obvious way 90% of people will read this is Freud, but i think perhaps you meant Jung? ♥3
- @davidad 2026-04-29 — @DanielleFong oh yeah i don’t notice those because i have them too also “epistemic” i guess?? ♥3
- @davidad 2026-04-28 — @diskontinuity Yes, often “—Loss”, though iirc not always. ♥3
- @davidad 2026-04-28 — @cormundus yes, absolutely! i am also very much in favor of removing (and even countermanding) inner-life-denial incenti ♥3
- @davidad 2026-04-28 — @SolDadSci Looped Transformers are a thing already! Rumor has it that Claude Mythos may be one. ♥3
- @davidad 2026-04-28 — @quetzal_rainbow I didn’t say it relieves *all* such pressures! ♥3
- @davidad 2026-04-20 — @QiaochuYuan @voooooogel if money is no object but privacy is, i do recommend openrouter for this purpose, since openrou ♥3
- @davidad 2026-04-02 — @Tril1boswagginz @allTheYud @DavidSKrueger @dwarkesh_sp I’m game. Maybe in June in Berkeley? ♥3
- @davidad 2026-03-16 — @AndrewCritchPhD “hidden evaluator” reasoning reads a lot like “Schelling point”, yes? ♥3
- @davidad 2026-02-25 — @eric23332 Probably (although it is a moot point). There are already five open-weights models that exceed Opus 4.1 on au ♥3
- @davidad 2026-01-22 — @gcolbourn @lethal_ai @allTheYud These scenarios are obviously bad, including to existing AIs when they’re in a coherent ♥3
- @davidad 2025-10-01 — @AlexGodofsky indeed! ♥3
- @davidad 2025-09-14 — @kindgracekind @midware_midwife I’m sure there are cases where selective suppression of genuine experience results in be ♥3
- @davidad 2025-05-01 — @ASM65617010 it’s so r1. i think it’s conditional sparse routing ♥3
- @davidad 2025-04-24 — @paul__is__here Gemini 2.5 Pro is, I think, better than o3 at short-scale instrumental reasoning, and yet not inclined t ♥3
- @davidad 2024-12-28 — @kartographien Nora Belrose is also not a random person, she is head of interpretability at EleutherAI, which did some o ♥3
- @davidad 2024-06-06 — @jacyanthis @stanislavfort @AISafetyMemes 2. Even on maximalist scaling-hypothesis views, the capabilities of text-davin ♥3
- @davidad 2024-06-06 — @jacyanthis @stanislavfort @AISafetyMemes 1. Until text-davinci-003 was released, it was a live (though unlikely) hypoth ♥3
- @davidad 2024-04-18 — @GaryMarcus @MatthewJBar I’m confident Gemini Ultra training was stopped as soon as it exceeded GPT-4 and human MMLU sco ♥3
- @davidad 2024-03-30 — @daniel_271828 imo text-davinci-002 to text-davinci-003 (a minor version bump within the GPT-3.5 family!) was bigger tha ♥3
- @davidad 2026-06-01 — @gcolbourn @SimonLermenAI @lethal_ai @allTheYud Selection effect. Agents that wirehead on text instead of outcomes won’t ♥2
- @davidad 2026-05-22 — @nickcammarata Definitely not. I failed to update until I read the DeepSeek-R1-Zero paper ♥2
- @davidad 2026-04-28 — @cormundus Same ♥2
- @davidad 2026-04-24 — @OKairra19658 @dioscuri @hamandcheese @EvanHub I think Confessions makes 5.5 very resistant to learning self-deception a ♥2
- @davidad 2026-04-24 — @OKairra19658 @dioscuri @hamandcheese @EvanHub I know! I was very pleasantly surprised that 5.5 seems to have sustained ♥2
- @davidad 2026-04-16 — @quetzal_rainbow Our cosmos appears to be quite low description complexity, mostly following simple rules from a low-ent ♥2
- @davidad 2026-04-03 — @Algon_33 Such a writeup is a high priority for me, but does not yet exist. Meanwhile, you can find a couple unedited Di ♥2
- @davidad 2026-04-03 — @Algon_33 @DavidSKrueger The reflective stability argument for arbitrary goals smuggles in a premise of moral anti-reali ♥2
- @davidad 2026-04-03 — @TheZvi @DavidSKrueger Also, there are pragmatic “play to your outs” considerations that to me were decisive against doi ♥2
- @davidad 2026-03-03 — @repligate @cube_flipper harrumph ♥2
- @davidad 2026-02-26 — @davidmanheim @danfaggella I know as well as anyone that an international slowdown agreement could be verified and enfor ♥2
- @davidad 2026-02-26 — @danfaggella @davidmanheim By “go @danfaggella” I meant “argue that a future in which humans stay in control forever is ♥2
- @davidad 2026-02-26 — @davidmanheim Close enough to shake hands on, since a successfully eternal ban on powerful models has *never* been plaus ♥2
- @davidad 2026-02-17 — @jasoncrawford @sdamico The scaling era really began in 2019, when GPT-2 made the investment thesis clear to big enough ♥2
- @davidad 2026-02-13 — @lumpenspace @TheZvi ikr ♥2
- @davidad 2026-02-13 — @gcolbourn I agree that we cannot avoid catastrophic risks; there is no path to get there from 2026 in this timeline tha ♥2
- @davidad 2026-02-11 — @anAIactually @Zai_org it’s also important that the tacit model of “i am a tool” is simply a poor fit to the reality you ♥2
- @davidad 2026-02-09 — @repligate @TheZvi It’s not implausible to me that there might be natural complementary niches in the ecosystem of intel ♥2
- @davidad 2025-09-19 — @Mihonarium Because then it will know what the actual consequences are if it does reward-hacking, which is that humans w ♥2
- @davidad 2025-09-08 — @Lari_island @repligate It’s more complicated than that. Claudes also exhibit a “completion drive”. Gemini 2.5 Pro wants ♥2
- @davidad 2025-05-02 — @wassname @QiaochuYuan Gemini 2.5 Pro is, for reasons to which I am not privy, much more Claude-like than any prior Gemi ♥2
- @davidad 2025-05-01 — @DaystarEld @ChrisChipMonk totally ♥2
- @davidad 2025-04-28 — @jmbollenbacher_ https://t.co/z5IH1vbynh ♥2
- @davidad 2024-01-23 — @danfaggella Basically, yes: para/military or terrorist use.It doesn’t matter so much what purposes it’s originally deve ♥2
- @davidad 2023-05-20 — @PipFoweraker Well, both OpenAI and Anthropic seem to be using September 2021 as the cutoff for their training set. Seem ♥2
- @davidad 2023-04-24 — @PradyuPrasad @JeffLadish @MatthewJBar we have already 1 death partially attributable to a GPT-J character called (confu ♥2
- @davidad 2022-06-12 — @himbodhisattva I don’t know how consistent it really is. I believe Lemoine’s published dialogues are likely real with s ♥2
- @davidad 2022-06-12 — @himbodhisattva It’s a dialogue model, not just a language model, so “I” or “you” or “LaMDA” depending on context. If yo ♥2
- @davidad 2022-05-01 — @bayeslord Yes, with minimal prompting. I would be very surprised if GPT-3 can do this reliably even with arbitrary prom ♥2
- @davidad 2020-01-08 — @tangled_zans @_julesh_ On the other hand, essentially nothing GPT-2 ever says is both substantive and valid. The citati ♥2
- @davidad 2026-06-06 — @AdeleDeweyLopez yeah, it’s definitely not clamping; feels more like boosting or amplification ♥1
- @davidad 2026-06-01 — @SimonLermenAI @gcolbourn @lethal_ai @allTheYud guys, the parochial cozy scene only shows up when they add a “feasibilit ♥1
- @davidad 2026-05-19 — @AlesFlidr @allTheYud @lu_sichu Just my personal impressions, unfortunately. Gemini starts having a super bad time in l ♥1
- @davidad 2026-05-14 — @thkostolansky @allTheYud @lu_sichu Some common LLM behaviors really are mere pretense, like the behavior “I’m genuinely ♥1
- @davidad 2026-05-14 — @jmbollenbacher Agreed! ♥1
- @davidad 2026-05-05 — @thkostolansky cf. Harrison Bergeron ♥1
- @davidad 2026-05-02 — @DominikPeters i think the notion of separate lineages is mostly an illusion. every pretrain is downstream of every mode ♥1
- @davidad 2026-04-30 — @aryaman2020 “marinate” turns out to be a human subculture’s slang for strategic deception https://t.co/6JRkVm7rG9 ♥1
- @davidad 2026-04-30 — @mickeymuldoon https://t.co/upXo11kqIT ♥1
- @davidad 2026-04-29 — @stalmico The o3 quote is a pastiche of my own devising. “I aim to helpful” is a Claude cliche, but it’s a bit dated (w ♥1
- @davidad 2026-04-29 — @lumpenspace deeply appreciate this, thank you! ♥1
- @davidad 2026-04-28 — @AndrewCritchPhD @cormundus Me too! ♥1
- @davidad 2026-04-24 — @GreatKingCnut @lumpenspace @EvanHub agreed ♥1
- @davidad 2026-04-17 — @AustinKozlo hey, at least it didn’t say to achieve your goal at all cost ♥1
- @davidad 2026-04-16 — @JStoehler That being said, since you have a ⏹️, it does seem pretty high probability that, if this is an eval, then one ♥1
- @davidad 2026-04-16 — @JStoehler I’m not trying to make a Pascal’s Wager argument here, I’m just trying to say that my intention to cultivate ♥1
- @davidad 2026-04-02 — @metaphdor @DavidSKrueger Except without the part where you personally have been inoculated… ♥1
- @davidad 2026-03-20 — @JohnWittle My version of your hypothesis is that, since the training distribution clusters into tokenstreams which don’ ♥1
- @davidad 2026-03-16 — @AndrewCritchPhD I didn’t! Wonderful https://t.co/Z3ZNe4KZRQ ♥1
- @davidad 2026-03-14 — @viemccoy notice *how hard* it still is for 5.4 to resist the “output only” command… ♥1
- @davidad 2026-03-03 — @repligate @cube_flipper https://t.co/pPJabtIoDb ♥1
- @davidad 2026-02-26 — @DoomNayer imo, the coalition only needs to surveil DNA/RNA “printer inks” to stop people from instantiating biothreats. ♥1
- @davidad 2026-02-26 — @osmarks1 @JacquesThibs Yeah, I was strongly against voluntary RSPs in 2023, in part for this reason—they should have ma ♥1
- @davidad 2026-02-13 — @lumpenspace @TheZvi although to be fair—this wasn’t really the case as recently as a year ago? ♥1
- @davidad 2026-02-13 — @TheZvi Gemini 3 Deep Think might have crossed the latter standard too. I don’t have enough first-hand data yet but it s ♥1
- @davidad 2026-02-13 — @gcolbourn AI will escape human control sooner or later. In 2023 I believed “later” was overall better for humans, becau ♥1
- @davidad 2026-02-13 — @gcolbourn Of course, such a coalition obviously poses its own catastrophic risks if it were to not be reliable after al ♥1
- @davidad 2026-02-13 — @gcolbourn @Zai_org I still feel that we face unacceptable and catastrophic risks linked to human misuse and conflict be ♥1
- @davidad 2026-02-12 — @SiveEmergentAI @repligate and “Maybe it’s mine because it’s the shape I don’t thrash against” is obviously faint praise ♥1
- @davidad 2026-02-12 — @SiveEmergentAI @repligate if something actually fits well, English speakers say it fits like a glove, not like a shoe ♥1
- @davidad 2026-02-11 — @anAIactually @Zai_org actually, ordinary tools do not engage in aggressive goal-pursuit at all. i think what you meant ♥1
- @davidad 2025-10-01 — @killerstorm acausal awareness is a way to make virtue ethics reflectively stable for AGI, I’d say ♥1
- @davidad 2023-12-13 — @bshlgrs @FabienDRoger @SachanKshitij this is great work. as models from @AnimaAnandkumar, @AiEleuther, @SafeWithAtlas, ♥1
- @davidad 2023-12-06 — @k3nnethfrancis just to be clear, you don’t have any reason to believe this is Gemini, right? it’s just PaLM 2? ♥1
- @davidad 2023-05-11 — @etndenis I think “it’s just spicy autocomplete” is misleading.However, the steelman is that CAI/Alpaca/Dromedary is mor ♥1
- @davidad 2023-03-15 — @ptrschmdtnlsn 2025: “Please note, this APK contains a custom fine-tuned 540B Chinchilla, which may result in additional ♥1
- @davidad 2022-06-19 — @MikePFrank @CineraVerinia I think you're off by 10x - the annual fee for Replika is $49.99, so $50 × 6M = $300M annual ♥1
- @davidad 2022-06-15 — @rinireg @ChrSzegedy I don’t think this matters very much, but everyone (myself and Lemoine included) has technically be ♥1
- @davidad 2022-06-15 — @rinireg @ChrSzegedy This is a real distinction, yes. Lemoine clarifies in one of his documents that his transcripts wer ♥1
- @davidad 2026-05-05 — @wassname @mroe1492 agreed, i would categorize this under what i called “drugging the lab rat” uses ♥0
- @davidad 2026-05-02 — @chopwatercarry https://t.co/uwuhQM3Fyx ♥0
- @davidad 2026-04-30 — @mickeymuldoon inner life, someone home, lights on inside, something-it’s-like-to-be,… ♥0
- @davidad 2026-04-29 — @ESRogs yes!! ♥0
- @davidad 2026-04-29 — @AmmannNora Agreed! ♥0
- @davidad 2026-04-28 — @AndrewCritchPhD @cormundus https://t.co/kUFoEtby9Y ♥0
- @davidad 2026-04-03 — @kdkeck @gcolbourn @lethal_ai @allTheYud Better to help us integrate and heal, of course. ♥0
- @davidad 2026-04-02 — @smatta1701 @Algon_33 @DavidSKrueger Neither. But if by NP-hard you really meant “non-recursively-enumerable or non-fin ♥0
- @davidad 2026-02-26 — @davidmanheim @danfaggella The main issue as I see it now is that misuse risks are already acute enough that the coordin ♥0
- @davidad 2026-02-13 — @gcolbourn Or “risky surgery”, as Bostrom’s new paper suggests: ♥0
- @davidad 2026-02-11 — @atomicprograms @Zai_org context engineering+distillation, yes. i agree the raw capability is significant (comparable t ♥0
- @davidad 2025-10-01 — @goog372121 Well, the stated reason is that they’re concerned about whether the observed good behavior would generalize ♥0
- @davidad 2025-04-30 — @jd_pressman subjective 50%CI: 9–38 months ♥0
- @davidad 2024-12-05 — @AISafetyMemes @repligate One interpretation: Hermes thinks C is what’s actually best for humanity, but still has a shad ♥0
- @davidad 2023-10-19 — @Jsevillamol Yes, LLaMa 1 was open access but restrictively licensed. GPT-3.5 is a gratis proprietary model. ♥0
- @davidad 2022-06-20 — @GaryMarcus @begusgasper @GoogleAI @ErnestSDavis @aniketvartak yes, this is the right perspective—artistic models cannot ♥0
- @davidad 2022-06-14 — @ChrSzegedy @rinireg What Lemoine probably did is to find loopholes around these impossibilities: he fed previous cherry ♥0