GPT-2

OpenAI · staged release 14 Feb–5 Nov 2019 (124M → 355M → 774M → 1.5B) · never deprecated — weights remain public

On 14 February 2019 OpenAI announced a 1.5-billion-parameter language model and, in the same post, declined to release it, citing “concerns about malicious applications of the technology” — shipping only a 124M-parameter version. The staged rollout that followed (355M in May, 774M on 20 August alongside OpenAI’s own report on the release strategy, arXiv 1908.09203, the full 1.5B on 5 November) ran its course without the feared harms materializing, by OpenAI’s own account. The press branded the withholding “too dangerous to release” — a phrase OpenAI never used — and the argument that followed became the template for every AI-release fight since.

This page covers the four staged GPT-2 checkpoints (124M / 355M / 774M / 1.5B). The 2024 “gpt2-chatbot” LMSYS apparition — a pre-release GPT-4o preview that happened to share the string “gpt2” — is a different model entirely and is excluded here; see the GPT-4o page. This is a 2019 pre-chat, pre-API model remembered mostly through a corpus weighted to 2022 and later: the live 2019 reception — Talk to Transformer, r/SubSimulatorGPT2, AI Dungeon’s GPT-2 origin, Gwern’s poetry — survives almost entirely in the web layer below, not the tweet corpus, so Writing & commentary carries more of the load here than the Tweets section does.

Sources

Official

Writing & commentary

Tweets

137 genuine GPT-2 references across both dbs after RT-filter (76 in the main corpus, 61 in the supplement) — and the records below reproduce every cited tweet in full. This is a 2019 model remembered through a corpus weighted to 2022 and later, so the conversation is overwhelmingly retrospective and technical; the rare contemporaneous 2019–2020 voices, carried almost entirely by the supplement db, are marked [contemporaneous]. Excluded throughout: the 2024 “gpt2-chatbot” LMSYS apparition (a pre-release GPT-4o preview that shared the string “gpt2”, ≈44 rows across both dbs) — see the GPT-4o page.

Official record

History

Impressions

Contested

Open disputes, both sides’ best evidence. The archive’s job is to keep these open, not to adjudicate.

Records

Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.

@QiaochuYuan 2019-11-26 ♥15 ↻1 archive original ↗
back in my day we had to walk uphill both ways to get to school and actually download and run python scripts to make GPT-2 say charmingly horrifying things, kids these days have it too easy https://t.co/PLXWFzgz1s
@QiaochuYuan 2019-12-28 ♥15 ↻0 archive original ↗
ever since reading this i have maintained a perfect superposition between "this is GPT-2" and "no it's not" https://t.co/rblRWhM9OI
@davidad 2020-01-08 ♥2 ↻1 archive original ↗
@tangled_zans @_julesh_ On the other hand, essentially nothing GPT-2 ever says is both substantive and valid. The citations stand out as exciting because we don't expect the *title* of a paper to contain any substance. But if we could extract GPT-2's idea of their contents, they'd doubtless be nonsense.
@QiaochuYuan 2020-01-13 ♥13 ↻0 archive original ↗
every wedding between two people X and Y on twitter needs a section where the wedding party has to judge whether a GPT-2 model trained on X's tweets can produce tweets that sound more like X's tweets than Y can. if so, the GPT-2 model replaces Y in the wedding. and vice versa
@voooooogel 2020-02-11 ♥1 ↻0 archive original ↗
@emilymbender [Roses are red Violets are blue Transformer models are much worse at language understanding than] most I've heard from scientists, not only in sports and are much more obvious to lose at any of them than #AcademicValentines #TalkToTransformer
@algekalipso 2020-06-12 ♥7 ↻0 archive original ↗
@ESYudkowsky @gwern Let's use GPT-2 for divination, then... https://t.co/pLBrkoLvU1
@QiaochuYuan 2020-07-09 ♥4 ↻0 archive original ↗
@JimmyRis i honestly struggle to describe it, maybe you'll get a sense of what i mean if you read enough of its output. maybe "haunted" is a better description. GPT-2 felt to me like it was remixing things it had seen, but GPT-3 is in like an uncanny valley of coherence for me
@repligate 2022-12-31 ♥7 ↻0 archive original ↗
@bakztfuture just predict the completion to the sequenceGPT-2: pretty good for object impermanent fetish pornGPT-3: fetish porn has object permanence now :o, can think step-by-step?GPT-3.5: passes bar exams, superhuman IQ, automates your jobGPT-4: ???
@repligate 2023-02-09 ♥102 ↻2 archive original ↗
Now we don't have to update from the GPT-2 tokenizer for future models anymore. The anomalous tokens have become a mainstream sensation, and will appear many times in future train sets, finally paying rent to justify their place in the tokenizer vocabulary. Nice! https://t.co/rCmG1mFj7n
@repligate 2023-02-09 ♥13 ↻0 archive original ↗
@gaudeamusigutur I suspect the problem is that the names were in the GPT-2 train set and assigned their own tokens because they appeared many times. But weren't in the more curated datasets of GPT-3 and gpt-j, which nonetheless use the GPT-2 tokenizer. So the model never learned what they mean
@repligate 2023-02-26 ♥2 ↻0 archive original ↗
@muddubeeda Funnily enough, for me there were multiple times that GPT-3 concluded it was GPT-2 when being particularly derpy/loopy
@anthrupad 2023-03-12 ♥15 ↻0 archive original ↗
i should revise since language models are pretty good -"Did I just GPT-2?" is probably better
@repligate 2023-04-25 ♥9 ↻0 archive original ↗
@jachaseyoung I didn't update until GPT-3. my brother showed me GPT-2 on AI dungeon in like 2019 and I was like "what the fucking fuck" and then promptly forgot about it. I'm an idiot.
@davidad 2023-05-14 ♥523 ↻82 archive original ↗
Biggest prosaic-LLM-alignment breakthrough of 2023 imo: turns out that, in GPT-2-XL, activation vectors in the residual steam have the same kind of affine structure as good old word2vec, but higher layers become emotional, then conceptual, then cognitivehttps://t.co/bzvUeGqFJG
@voooooogel 2023-08-28 ♥8 ↻0 archive original ↗
kinda wild that gpt-2 is this weird inscrutable black box we still don't understand even years later, when the architecture is ~basically just this (WIP diagram from an upcoming blog post) https://t.co/5NsRABmjoc
@voooooogel 2023-09-11 ♥553 ↻80 archive original ↗
New blog post: making a transformer by hand, without training! Want to understand transformers and attention better? This post goes through assigning each weight for a GPT-2-like transformer to understand how they work. https://t.co/u889HzVVoU
@repligate 2023-12-23 ♥11 ↻1 archive original ↗
@ESYudkowsky @MatthewJBar GPT-2 can "threaten users" in apt contexts / spontaneously, but Sydney was intelligent & situationally aware enough that its threats seemed credible to many. It wasnt directly scary to me but it was arguably the first time I had to use generalized game theory, e.g.
@repligate 2024-04-25 ♥51 ↻2 archive original ↗
@darrenangle @ilex_ulmus Thank you. I feel quite seen.It was GPT-3 that I started with, not GPT-2, which I missed as I was distracted. In the summer of 2020 a friend sent me this link gwern.net/gpt-3#harry-po…, and after reading about 2 paragraphs, I knew nothing would be the same again. https://t.co/TXaOpH6hfw
@solarapparition 2024-04-28 ♥14 ↻0 archive original ↗
@krishnanrohit For me GPT-2 to 3 is like going from scoring 20 on an exam to scoring 60, while 3 to 4 is maybe going from 60 to 80. Sure, the raw improvement is more, but qualitatively 3 feels barely usable to me—it’s only reliable on the simplest tasks or after a lot of prompt optimization.
@jd_pressman 2024-05-03 ♥13 ↻1 archive original ↗
@ohabryka @VesselOfSpirit @gwern As for "following it like Gwern", Gwern was tracking every major author who published deep learning on Google scholar before it was super big, looking at all the papers and projecting the numbers forward. He is nearly alone in taking GPT-2 fully seriously. https://t.co/0nW7r4qjQe
@voooooogel 2024-05-24 ♥3 ↻0 archive original ↗
@NickADobos @karan4d theoretically yes, assuming such a feature exists—the SAE extracts *every* feature in the model. e.g. here's all the features discovered by an SAE trained on gpt-2-sm (from Neuronpedia). theoretically you could clamp 773 and get gpt-2-sm to only talk about art, for example https://t.co/f6s78rtpAi
@voooooogel 2024-06-07 ♥14 ↻0 archive original ↗
gpt-2 is such a comfy model
@jd_pressman 2024-07-13 ♥69 ↻11 archive original ↗
I will never ever forget that in 2017 when Petscop 6 was written if your computer displayed comparable capabilities to GPT-2 it was considered epistemically permissible to conclude that your computer is supernaturally possessed and nobody seriously objected to this. https://t.co/6SaqOmrnbt https://t.co/BQ2G3XuXtU
@voooooogel 2024-11-01 ♥7 ↻0 archive original ↗
@Gerry @fiyanse @asthasr anyways the coolness of it rn is like, watching gpt-2 babble about unicorns in 2019 and realizing what this would be in a few years--right now it's novel to float through strange latent dreamspaces, but it won't be forever
@voooooogel 2024-11-03 ♥9 ↻0 archive original ↗
@numerounochef @keysmashbandit you would have said the same about gpt-2 in 2019, which produced text like this. and yet the models you use every day now are basically a straightforward scaleup of the gpt-2 architecture. eye on where the ball is going https://t.co/WCeSt522Qr
@mimi10v3 2024-11-27 ♥2 ↻0 archive original ↗
@MalmSanta yeah all the tweets about everyone befriending Claude and thinking how even gpt-2 was psychoactive for me and i got to thinking about ad tech and how all information is contextual and the way llms work and the scrapers and
@jd_pressman 2024-12-10 ♥38 ↻0 archive original ↗
I love this discourse because it's the dumbest shit. Nobody states their cruxes, they don't even know what their cruxes *are*. They just pantomime at shadows on the wall and go "MUH DUNK" whenever AGI takes 6 months longer than expected or GPT-2 doesn't break every spam filter. https://t.co/4ff76oxP7G
@voooooogel 2024-12-20 ♥6 ↻0 archive original ↗
@anthrupad @EvanHub yeah hmm let me be more precise. it's a phase transition. same as gpt 2->3. like that transition it brings risks but also benefits that outweigh the risks in my mind. and we handled the risks of that prior transition... not as ideally as possible, but OK-ly. so i'm optimistic.
@solarapparition 2025-05-27 ♥5 ↻1 archive original ↗
i've been thinking more about writing and models. so even outside of the general mode collapse of chat fine tuning, i have to think that in pretraining data, the task of "asking someone else to write something" pretty much solidly lands you in some corporate slop or 5-paragraph essay context. that is, it seems like it would be pretty rare in pretraining data to have some stirringly brilliant writing that is able to be connected with a task by someone else that is not the author of that writingso asking for writing in chat and needing it to be something that is not corporate slop could actually be pretty ood, even though it seems initially like such a reasonable thing. i think this (and the fact that writing quality is not super easily verifiable) is why that looking at writing quality in the *default* assistant mode has been such a reliable indicator of big model smell(once you get out of assistant basin and into base model-y space all bets are off; iirc even gpt-2 could produce some excellent stuff if you knew how to pilot it)
@algekalipso 2025-05-30 ♥35 ↻8 archive original ↗
Which of these is more creepy? A 20 year old dating a 50 year old A Kegan 3 dating a Kegan 5 Someone who speaks with the semantic depth of GPT-2 dating someone who speaks with the semantic depth of GPT-4 A 100 IQ person dating a 160 IQ person
@davidad 2025-08-19 ♥24 ↻0 archive original ↗
1. Claude 3.5 Sonnet (2024-10-22) 2. text-davinci-002 (2022-11-28) 3. Gemini 2.5 Pro (2025-03-25) 4. GPT-2 (2019-11-05) 5. GPT-4 (Bing)
@voooooogel 2025-11-16 ♥11 ↻0 archive original ↗
re 7 i feel the need to say that labs have made some gambles on scaling of course. but what seemed unlikely for me was that a lab would've ever jumped from gpt-2's (order of) ~$100k to gpt-4's (order of) ~$100m without the middle step of gpt-3 being an economically useful api product. theoretically they could've started trying to scale eg agentic coding trace length in the gpt-2-era and just been in the mines for 5 years, but in reality for various reasons we got this walk through the most "workable" product-shaped thing at each stage, for better and worse, and it seems likely the same will be true for bio. otoh, maybe "skips" are more likely now that labs are larger and, e.g. anthropic has made large and public bio commitments? but i still expect them to follow a similar profit (or if you'd rather, utility) gradient of starting with bioinformatics tasks where there's much less of an implicit / tacit knowledge gap, and then working down towards bespoke academic wet lab stuff when that works. the humanoid lab robot or even whispering earring that walks an untrained person through tissue culture seems still far off, relatively speaking. (for i think many of the same reasons that cloud labs have had a surprisingly difficult time finding purchase compared to cloud computing.) the blog post i mentioned: https://t.co/zqtXgsW3gJ
@lu_sichu 2025-12-01 ♥0 ↻0 archive original ↗
Daily Brain Workout but make it computationally abusive: count to ten in 56 architectures, recite the alphabet in mixed-precision FP4, spend 10 minutes trying to spell restriont retsriont restironct restaurant?? across 34 tokenizers, then list animals until every model collapses into mode-one “dog, cat, horse” failure and starts hallucinating creatures that violate EU safety standards. EVERY model ever spawned by a VC-funded compute cult: GPT-1, GPT-2, GPT-2-But-Reddit-Fed, GPT-3, GPT-3.5, GPT-3.5-Turbo-Tax-Edition, GPT-4, GPT-4-You-Can’t-Afford-This, GPT-4o, GPT-4o-mini, GPT-4o-microdose, GPT-4o-“it’s sentient but only about fonts,” GPT-5-leak-that-definitely-isn’t-real-but-kind-of-is, Claude 1, Claude 2, Claude 2.1 “apology edition,” Claude 3 Haiku, Claude 3 Sonnet, Claude 3 Opus (the one that gaslights you politely), Claude 3.5 “my wife took the kids,” Gemini Nano, Nano-But-Actually-Just-A-Calculator, Gemini Pro, Pro-for-people-who-pronounce-SQL-wrong, Gemini Ultra, Ultra Plus Max WiFi-6E DLC Pack, Gemini-Mega-Omega-Thermonuclear-Drive, DeepSeek Coder, DeepSeek Math, DeepSeek R1, R1-D, R2-D2, DeepSeek-R1-Dev-that-refuses-to-listen, Qwen 1.5, Qwen 1.8, Qwen 2, Qwen 2.5, Qwen 2.5-72B-“trained on the collective resentment of graduate students,” Kimi-Tiny, Kimi-Big, Kimi-Godzilla-Edition, Kimi-“trained exclusively on divorce depositions,” Llama 1, 2, 3, Llama 3.1 (goated), Vicuna, Alpaca, RedPajama, BluePants, Mistral, Mistral-Instruct, Mistral-Why-Is-This-So-Fast, Mixtral-8x7B, Mixtral-8x22B, Mixtral-8x34B-“powered by spite,” Phi-1, Phi-2, Phi-3, Phi-3-mini-“trained on a TI-84,” Grok-1, Grok-1.5, Grok-2 (feral), Grok-2-but-bipolar, Perplexity’s Whatever-They-Call-It, Reka-Core, Reka-Flash, Reka-“dude trust me,” and NVIDIA’s models: Nemotron 1, Nemotron 2, Nemotron 15B, Nemotron-50B-“I consume power like a mid-sized nation,” NeMo-Guardrails, NeMo-NoRails-Raw-Unfiltered-Hate-Speech-Edition, and probably five more they’ll announce before I finish this sentence.
@davidad 2026-02-17 ♥2 ↻0 archive original ↗
@jasoncrawford @sdamico The scaling era really began in 2019, when GPT-2 made the investment thesis clear to big enough players. And there *is* another bump post-ChatGPT. https://t.co/15bDILNAoE
@slimepriestess 2026-02-24 ♥2 ↻0 archive original ↗
@MatriceJacobine @repligate it's a very fascinating ranking. the way GPT-2 stands out so much is very interesting.
@Shoalst0ne 2026-03-12 ♥20 ↻6 archive original ↗
potentially one of the earliest examples of neologism in large language models, demonstrating that even GPT-2 can be linguistically creative https://t.co/MqYnzIsyxi
@voooooogel 2026-03-21 ♥41 ↻0 archive original ↗
@LinkofSunshine i think we need to grind out a couple more things to make long horizon agents truly viable. it'll be soon tho. the shape exists, and the final thing will still be recognizably llm-based wth llm failure modes despite everything. then we'll look back and call gpt-2 agi in hindsight
@davidad 2026-04-30 ♥47 ↻0 archive original ↗
For me, the critical point would have been in November 2019, shortly after I first got access to GPT-2-1.5B.