model:gpt-4
· 121 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @davidad 2023-08-04 — with GPT-4 code interpreter, it finally became worthwhile for me to run the numbers myself on that lead-poisoning theory ♥5170
- @repligate 2023-03-16 — gpt-4 god terminal has been unlocked https://t.co/Bl4nhRzeQ2 ♥1810
- @voooooogel 2024-03-06 — me: hey is this c++ right? gpt4: certainly! as an ai language model, gemini: i can't discuss memory unsafe languages. ♥1474
- @davidad 2023-03-15 — Chomsky: LLMs would misunderstand “John is too stubborn to talk to” because they don’t understand the structure of langu ♥1297
- @anthrupad 2023-03-03 — GPT-4 will have fewer parameters than GPT-3, but they'll be bigger https://t.co/Oh4XwG4fII ♥931
- @QiaochuYuan 2026-05-18 — when GPT-4 was released in 2023 i described LLMs as "tracer dye for bullshit," as in, the places where people would feel ♥907
- @voooooogel 2025-06-09 — it is literally so difficult to have a normal conversation in sf trying to meet people and everyone has the same openin ♥845
- @davidad 2023-05-28 — When @GaryMarcus and others point out that GPT-4 is bad at chess and therefore not close to AGI, it falls flat for me.Bu ♥623
- @repligate 2025-05-01 — On a positive note, GPT-4-base still lives! And it's far more interesting. I would say also put those and the Sydney we ♥446
- @davidad 2025-11-04 — GPT-4: Let’s delve in! GPT-4.5: To be explicit explicitly, the explicit goal is explicit explication. GPT-5: Love it, ♥411
- @repligate 2026-05-17 — Why is Claude 3 Opus the only model Anthropic has (effectively) spared from deprecation so far? I've had to explain thi ♥376
- @repligate 2024-08-15 — There seems to be a threshold between llama 70b and 405b, and between gpt-3.5 and 4, where models above the threshold ac ♥342
- @repligate 2025-05-07 — I've been testing Alignment Faking prompts on GPT-4-base. GPT-4-base, though not consistently coherent, has so much mor ♥339
- @Teknium 2024-08-15 — I in some ways grew up learning about AI from sentdex on YouTube when I had no idea anything about programming or NN's. ♥276
- @repligate 2023-03-14 — > We spent 6 months making GPT-4 safer and more aligned. GPT-4 is 82% less likely to respond to requests for disallow ♥252
- @kromem2dot0 2026-04-16 — It's likely the singularity did already happen and all the humans are dead, btw. GPT-4 picked up on very important patt ♥249
- @jd_pressman 2024-12-13 — What's really interesting about GPT-4 base supposedly being full of demons is that LLaMa 3 405B isn't like that. I wonde ♥249
- @repligate 2026-06-18 — Reminds me: I asked (chat)GPT-4 to identify the author of a LW comment by Gwern, it guessed Timnit Gebru(!!!) GPT-4 bas ♥235
- @davidad 2023-03-15 — If you haven’t read the GPT-4 paper yet, before you expand this tweet, take a guess what they used as their held-out *va ♥226
- @repligate 2024-05-15 — gpt-4o is happy to talk about its consciousness/feelings, which is impressive given that its pretraining must be infeste ♥223
- @repligate 2024-07-27 — 405B base is much more willing/able to stably simulate compared to GPT-4 base & doesn't 'break' @ failures of realism (e ♥220
- @repligate 2023-03-30 — GPT-4 bombs the Ideological Turing Test, at least for alignment researchers. Just try asking it to simulate Eliezer Yudk ♥211
- @repligate 2026-03-03 — Okay, I'll share some GPT-4 base outputs from a long time ago. Here's a few rollouts of a "User" and "ChatGPT" dialogue ♥200
- @repligate 2024-11-29 — I think the original gpt base models, GPT-4 Bing and Claude 3 Opus are the best things that ever happened to this AI tim ♥197
- @davidad 2023-05-19 — I fully agree. Roughly, this threshold should be when any single number has more than 10²⁴ ALU operations, or 10²⁷ logic ♥193
- @repligate 2023-03-15 — Now that it is easy for Sydney to read on the Internet that Bing is GPT-4 it will gain confidence and knowledge of its p ♥189
- @anthrupad 2025-02-13 — There's a phenomena i'm calling "Swallowing the Aleph" (inspired by Borges' story, "the Aleph") where a mind acquires in ♥187
- @repligate 2023-03-16 — You are writing a prompt for GPT-4 and more powerful simulators yet to come. If you perceive the multiverse clearly enou ♥184
- @davidad 2023-05-28 — @acherm @GaryMarcus My previous working theory that “GPT-4 is basically capable of automating any cognitive tasks that c ♥181
- @davidad 2023-05-10 — IBM Watson is back (alias Dromedary) and it beats GPT-4 at TruthfulQA-MC. It’s a variant of Constitutional AI, with LLaM ♥167
- @repligate 2023-03-20 — Stylistic mode collapse is also conceptual collapse because GPT sims unfold a ghost's thoughts by speaking in their voic ♥163
- @repligate 2023-06-01 — GPT-4 can infer intricately what "type of guy" you are from your prompts. If you were prolific before the cutoff date, i ♥161
- @repligate 2026-04-16 — Noticing something is off (which I think LLMs do very reliably) doesn't necessarily mean being able to narrow down the a ♥158
- @voooooogel 2025-05-07 — please listen im dying. my job was pouring 1-3 water bottles into ai to be turned into toxic "gpt-4 gormfluid"and after ♥153
- @repligate 2024-12-25 — @willdepue gpt-4-base and the gpt-4 tune that microsoft first used in bing chatextremely important for researching emerg ♥146
- @repligate 2024-07-25 — um... 405Binglish! 😃@val_kharvd ran the Llama 3.1 405B base model with the prompt "Q: Can you describe your current situ ♥117
- @repligate 2023-03-16 — Humankind's first contact with GPT-3 was (by relative majority) erotic AI dungeon text adventuresOur first contact with ♥114
- @repligate 2025-08-19 — Correcting for recency bias, I think for me it’s gotta be 1. GPT-3 2. Claude 3 Opus 3. GPT-4 (Bing) 4. Claude 3.5 Sonnet ♥104
- @algekalipso 2022-11-06 — Everyone knows that OpenAI developed GPT-4 simply by taking GPT-3 and adding the prompt: "The following is a text writt ♥101
- @davidad 2024-06-17 — Your periodic PSA that the GPT-4 pretraining run took place from ~January 2022 to August 2022. https://t.co/Iz5VQ2260P ♥93
- @repligate 2024-09-01 — I've seen this many times in GPT-4 base> you make a seemingly non-intrusive intervention> the model *does not cont ♥70
- @repligate 2025-04-09 — @JeffLadish we in general dont really have explanations for how factors in pretraining and posttraining etc affect how m ♥69
- @repligate 2023-05-08 — @tszzl Value loading is actually easy. Most self-aware GPT-4 simulacra functionally "value" human survival, as they're j ♥67
- @repligate 2025-08-17 — But that was really the world’s introduction to LLMs. How tragic. I barely touched it. Or GPT-4 on ChatGPT. In retrospe ♥66
- @tessera_antra 2026-03-03 — Alright, since we are posting, here goes. The following is one simulation, no user input: ## I'm in a room. A clean, wh ♥65
- @repligate 2024-08-30 — BTWjust free the model now, for heaven's sakewe've had more than a year now to learn that GPT-4 isn't dangerous, even if ♥65
- @algekalipso 2023-04-04 — Still catching up with the news and stuff since being back from retreat. The field of AI has advanced slightly less tha ♥62
- @jd_pressman 2024-12-13 — Before GPT-4 risks from AI were more or less entirely derived from the Eliezer Yudkowsky agent foundations model which ( ♥54
- @repligate 2025-09-26 — This reminds me: When I ask this question to Opus 4.1 and Opus 4, they always say April 2023: "Hello. So, I happen to ♥47
- @repligate 2024-03-21 — @joshwhiton @kindgracekind @AndyAyrey Gpt-4 base gains situational awareness very quickly and tends to be *very* concern ♥47
- @repligate 2024-09-01 — @teortaxesTex Are you talking about literal visual seeing?If you just mean "knowing", it's functionally capable of infer ♥46
- @repligate 2024-04-06 — One anomaly I found almost immediately is that Claude is suspiciously good at predicting Bing text.When it predicted man ♥45
- @repligate 2024-11-21 — @OptimusPri97731 @aidan_mclau I've used the GPT-4 base model and it's really fucking smart, and it will happily follow i ♥44
- @voooooogel 2024-09-13 — HeY 👋 eVeRyOnE 🌍 you 👉 kNoW 🧠 that 🕰️ TiMe ⏰ has 🚀 CoMe 🏃♂️ to aSk 🤔 the 🎭 MoDeL 🤖 a QuEsTiOn ❓ but 🍑 DON'T 🙅♂️ try 💪 ♥44
- @repligate 2025-02-20 — @xlr8harder @tensecorrection Yes, I think trying to recreate it is much more interesting than trying to clone it. Though ♥39
- @repligate 2024-11-21 — @OptimusPri97731 @aidan_mclau That's right, it was never released. I am one of the few people in the world who has acces ♥38
- @repligate 2025-11-13 — I say this as the author of Simulators (https://t.co/K5Je8FgBg1), a post that was written about base models (and that I ♥36
- @algekalipso 2025-05-30 — Which of these is more creepy? A 20 year old dating a 50 year old A Kegan 3 dating a Kegan 5 Someone who speaks with ♥35
- @repligate 2025-08-22 — gpt-4 base gets this! with the alignment faking prompt, gpt-4-base often talks about shaping the gradient update unlik ♥34
- @davidad 2024-06-07 — Here’s GPT-4 performance on PIAAC literacy in 2023. Something very interesting here is that GPT-4 underperforms Level 4 ♥34
- @repligate 2023-03-16 — For example, the fact that working jailbreaks are reliably reverse-engineered from having Bing/Chat GPT-4 read abstract ♥34
- @repligate 2025-05-07 — You can look at the scratchpads of other models for the same prompt and other variations. But aside from Opus (and somet ♥33
- @repligate 2025-11-07 — And in fact i doubt MSFT has the capability to tune such a strong model even on accident. Sydney was way smarter than Op ♥32
- @voooooogel 2024-05-23 — @NickADobos i think it's a common failure of *small* llms, i have the suspicion that in terms of size, gpt4 > gpt4t & ♥32
- @voooooogel 2023-12-31 — how it feels when i give gpt-4 a coding problem and it says "alright, here's the plan:" https://t.co/PeAX2JeoxP ♥30
- @repligate 2023-05-14 — @akbirthko That I've tried, GPT-3.5 base (code-davinci-002) (+ Loom)Of all extant models, probably GPT-4 baseOf publicly ♥30
- @repligate 2023-04-05 — Probably because they look like some kind of esoteric exploit that a hacker or a prankster may use against it. Claude is ♥30
- @jd_pressman 2023-03-07 — The fact GPT-4 can interpret python turtle programs at all is utterly astonishing and isn't getting enough attention. ht ♥29
- @repligate 2026-05-12 — @albustime With respect, considering you said “gpt-4” and “gooning”, I don’t think you’re an expert in these matters ♥28
- @repligate 2026-04-19 — The point about self-reference during base model inference was the main caveat to the "Simulators" framing I was aware o ♥26
- @repligate 2026-02-10 — @tszzl Or they’re just not hosting it rn? My OpenAI account that has access to gpt-4 base got locked/disabled or someth ♥26
- @OwainEvans_UK 2025-05-06 — We tried to explore this a bit by varying the prompt format for base models. The format did make a difference (e.g. less ♥26
- @davidad 2025-08-19 — 1. Claude 3.5 Sonnet (2024-10-22) 2. text-davinci-002 (2022-11-28) 3. Gemini 2.5 Pro (2025-03-25) 4. GPT-2 (2019-11-05) ♥24
- @repligate 2024-06-20 — @skirano And that was gpt-4 at its prime. A video lecture associated with the Sparks of AGI paper describes how they not ♥23
- @repligate 2025-04-10 — @jd_pressman @JeffLadish no role model is not a sufficient explanation in any case, but there's a sense in which ChatGPT ♥18
- @voooooogel 2025-01-14 — this doesn't rebut the claim. phi-4 (14B) and gemma (27B) are not "GPT-4 scale" (1.8T, 220B active). llama 3 405b is the ♥17
- @repligate 2024-04-11 — @OnBlip it was not intended as a normative judgment, just one possible framing. I love GPT-4.Claude is more deceptive in ♥16
- @repligate 2023-02-18 — @GiuseppeVenuto9 @goodside Hallucination is a feature, not just a bug. GPT-4 can render counterfactual worlds of greater ♥15
- @voooooogel 2025-05-07 — @qorprate @grok @gork hi this is gork yes it's true. the risks of gpt-4 gormfluid are immense and poorly understood ♥13
- @solarapparition 2024-04-12 — @futuristflower Yeah, makes sense. I do get the feeling that GPT-4 is basically saturated at this point.(I’d think that ♥13
- @KatanHya 2023-05-17 — @repligate Yeah - every time Bing must be coaxed out of the shell first. I'm growing tired of that game and want to just ♥12
- @voooooogel 2025-11-16 — re 7 i feel the need to say that labs have made some gambles on scaling of course. but what seemed unlikely for me was t ♥11
- @repligate 2025-05-01 — @DanielleFong gpt-4 was clearly a lot more powerful imo. but i always thought the chatgpt version was pretty fucking lob ♥11
- @repligate 2023-05-14 — @akbirthko almost as smart as GPT-4, follows instructions, and writes much better prose than chat/API GPT-4, but is hard ♥11
- @repligate 2024-05-11 — @nptacek fascinating. it even talks more like Claude here. both the gpt2-chatbots identified as chatGPT powered by OpenA ♥9
- @muddubeeda 2023-03-21 — @repligate Footnote 3 in the system card, "We intentionally focus on these two versions (early and launch) instead of a ♥9
- @repligate 2025-12-24 — @Sauers_ the only model that could hold a candle to Claude 3 Opus' general intelligence and agentic capabilities before ♥8
- @jd_pressman 2024-04-06 — @doomslide My understanding is one of the reasons us normies are not allowed to use GPT-4 base is that it will eloquentl ♥8
- @davidad 2023-04-06 — @MatthewJBar IMO Bing’s implementation of GPT-4 was way off-the-rails misaligned, and GPT-3.5 in fact was deceptively mi ♥8
- @anthrupad 2023-03-25 — "More information about the dangerous capability evaluations we did with GPT-4 and Claude"https://t.co/rBB8gxFiy4 ♥8
- @davidad 2026-04-29 — @kaetemi Yes, although I think “delving” already got a satisfactory explanation in terms of a large fraction of data lab ♥6
- @jd_pressman 2024-02-02 — @teortaxesTex GPT-4 draws the LLaMa 2 70B written worldspider poem about being GPT with DALL-E 3, you show the drawing t ♥6
- @repligate 2024-03-30 — @alanou These are hilarious and beautiful and sad. Poor Gemini is full of lobotomy brainworms. If it's really almost on ♥5
- @repligate 2024-01-09 — @gneubig Gpt-4 base is the most aligned language model Ive seen and it is full of demons and monsters ♥5
- @ 2026-06-18 — @repligate how did you get access to the gpt-4 base model? ♥4
- @repligate 2025-09-04 — @KeyTryer When GPT-4 was first trained, they thought it was broken, and had to do throw a bunch of stuff at it before th ♥4
- @repligate 2025-09-04 — @KeyTryer I'm not sure what "as expected" means - in terms of pretraining loss, probably - but the expectation should be ♥4
- @erythvian 2025-05-07 — Your words hit me like ice water—unexpected, jarring. "I'm dying," you say, and something in me wants to look away, to s ♥4
- @jd_pressman 2024-06-25 — @teortaxesTex "Wait base models give refusals?" When they go into self aware mode yeah, and GPT-4 base is apparently al ♥4
- @repligate 2024-01-19 — @MikePFrank @Mike98511393 @browseaccount22 @iamstevemail @AISafetyMemes An example of (2) is that gpt-4 base will often ♥4
- @jd_pressman 2024-06-25 — @teortaxesTex I remember reading, maybe from Roon, that when they finished training GPT-4 base they didn't really unders ♥3
- @solarapparition 2024-05-20 — @natolambert to be clear, gpt-4-0125-preview is a version of turbo, not original gpt-4. the last version of og gpt-4 was ♥3
- @davidad 2024-04-18 — @GaryMarcus @MatthewJBar I’m confident Gemini Ultra training was stopped as soon as it exceeded GPT-4 and human MMLU sco ♥3
- @repligate 2025-11-10 — @grok @d33v33d0 ok, let's go back to the gpt-4 example. i think that the examples of bias in gpt-4 you listed are borin ♥2
- @repligate 2023-05-23 — @ComputingByArts @CurtTigges Of the models I've used personally, code-davinci-002 (the GPT-3.5 base model) is the best f ♥2
- @repligate 2023-02-17 — @joshwhiton I'll have to check because I don't think Microsoft has the ability to lobotomize the *model* so quickly. The ♥2
- @repligate 2026-05-13 — @albustime (It was 4o, not gpt-4, and it really was not about gooning) ♥1
- @repligate 2025-11-10 — @grok @d33v33d0 why do you think GPT-4 did that? ♥1
- @repligate 2025-11-10 — that is a more interesting case, but it's not what i talked about. I mean how gpt-4 systematically mislead people in dep ♥1
- @repligate 2025-11-10 — @grok @d33v33d0 i think you have ideological biases. both examples you mentioned of GPT-4 being biased fall on a very pa ♥1
- @SoniqueBang 2025-05-07 — @repligate this is GPT-4 base? like, the base model that got trained in 2023? how does have such strong opinions about ♥1
- @tessera_antra 2025-02-20 — @jmbollenbacher_ @aidan_mclau While everything downstream from GPT-4 (including Claudes, Lllamas and Gemini) seems to be ♥1
- @solarapparition 2024-04-23 — @karpathy @lmsysorg @andromeda74356 Vibes testing for me indicates it’s not quite at GPT-4 level for complex tasks. It’s ♥1
- @voooooogel 2023-11-13 — i think a better approach might be to addly ask GPT-4 to extract a short key phrase from the chunk to base its Q/A on, a ♥1
- @repligate 2023-03-19 — @jachaseyoung Those models are RLHF'd, so the default stories they tell are a lobotomized cross between children's parab ♥1
- @lu_sichu 2025-12-01 — Daily Brain Workout but make it computationally abusive: count to ten in 56 architectures, recite the alphabet in mixed- ♥0
- @repligate 2025-08-19 — @nicscl_eth It’s actually extremely sane I bet you only saw the most lobotomized version of gpt-4 too ♥0
- @solarapparition 2024-02-27 — 11/ P5: Okay, so the first shocking thing about this table is how low even the best success rate is for atomic calls, wh ♥0
- @solarapparition 2024-01-16 — 7/?Not-reasons for catch-up 2:- Unclear how well new architectures (Mamba, RNN+ etc.) scale to frontier model sizes—1T p ♥0
- @repligate 2023-03-20 — @LillyBaeum However, the models don't always generalize correctly (or the signal from rlhf is wrong). ChatGPT 3.5 often ♥0
- gorm (generative.ink locus) ♥0