on:gpt-2
· 38 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @voooooogel 2023-09-11 — New blog post: making a transformer by hand, without training! Want to understand transformers and attention better? Thi ♥553
- @davidad 2023-05-14 — Biggest prosaic-LLM-alignment breakthrough of 2023 imo: turns out that, in GPT-2-XL, activation vectors in the residual ♥523
- @repligate 2023-02-09 — Now we don't have to update from the GPT-2 tokenizer for future models anymore. The anomalous tokens have become a mains ♥102
- @jd_pressman 2024-07-13 — I will never ever forget that in 2017 when Petscop 6 was written if your computer displayed comparable capabilities to G ♥69
- @repligate 2024-04-25 — @darrenangle @ilex_ulmus Thank you. I feel quite seen.It was GPT-3 that I started with, not GPT-2, which I missed as I w ♥51
- @davidad 2026-04-30 — For me, the critical point would have been in November 2019, shortly after I first got access to GPT-2-1.5B. ♥47
- @voooooogel 2026-03-21 — @LinkofSunshine i think we need to grind out a couple more things to make long horizon agents truly viable. it'll be soo ♥41
- @jd_pressman 2024-12-10 — I love this discourse because it's the dumbest shit. Nobody states their cruxes, they don't even know what their cruxes ♥38
- @algekalipso 2025-05-30 — Which of these is more creepy? A 20 year old dating a 50 year old A Kegan 3 dating a Kegan 5 Someone who speaks with ♥35
- @davidad 2025-08-19 — 1. Claude 3.5 Sonnet (2024-10-22) 2. text-davinci-002 (2022-11-28) 3. Gemini 2.5 Pro (2025-03-25) 4. GPT-2 (2019-11-05) ♥24
- @Shoalst0ne 2026-03-12 — potentially one of the earliest examples of neologism in large language models, demonstrating that even GPT-2 can be lin ♥20
- @anthrupad 2023-03-12 — i should revise since language models are pretty good -"Did I just GPT-2?" is probably better ♥15
- @QiaochuYuan 2019-12-28 — ever since reading this i have maintained a perfect superposition between "this is GPT-2" and "no it's not" https://t.c ♥15
- @QiaochuYuan 2019-11-26 — back in my day we had to walk uphill both ways to get to school and actually download and run python scripts to make GPT ♥15
- @voooooogel 2024-06-07 — gpt-2 is such a comfy model ♥14
- @solarapparition 2024-04-28 — @krishnanrohit For me GPT-2 to 3 is like going from scoring 20 on an exam to scoring 60, while 3 to 4 is maybe going fro ♥14
- @jd_pressman 2024-05-03 — @ohabryka @VesselOfSpirit @gwern As for "following it like Gwern", Gwern was tracking every major author who published d ♥13
- @repligate 2023-02-09 — @gaudeamusigutur I suspect the problem is that the names were in the GPT-2 train set and assigned their own tokens becau ♥13
- @QiaochuYuan 2020-01-13 — every wedding between two people X and Y on twitter needs a section where the wedding party has to judge whether a GPT-2 ♥13
- @voooooogel 2025-11-16 — re 7 i feel the need to say that labs have made some gambles on scaling of course. but what seemed unlikely for me was t ♥11
- @repligate 2023-12-23 — @ESYudkowsky @MatthewJBar GPT-2 can "threaten users" in apt contexts / spontaneously, but Sydney was intelligent & s ♥11
- @voooooogel 2024-11-03 — @numerounochef @keysmashbandit you would have said the same about gpt-2 in 2019, which produced text like this. and yet ♥9
- @repligate 2023-04-25 — @jachaseyoung I didn't update until GPT-3. my brother showed me GPT-2 on AI dungeon in like 2019 and I was like "what th ♥9
- @voooooogel 2023-08-28 — kinda wild that gpt-2 is this weird inscrutable black box we still don't understand even years later, when the architect ♥8
- @voooooogel 2024-11-01 — @Gerry @fiyanse @asthasr anyways the coolness of it rn is like, watching gpt-2 babble about unicorns in 2019 and realizi ♥7
- @repligate 2022-12-31 — @bakztfuture just predict the completion to the sequenceGPT-2: pretty good for object impermanent fetish pornGPT-3: feti ♥7
- @algekalipso 2020-06-12 — @ESYudkowsky @gwern Let's use GPT-2 for divination, then... https://t.co/pLBrkoLvU1 ♥7
- @voooooogel 2024-12-20 — @anthrupad @EvanHub yeah hmm let me be more precise. it's a phase transition. same as gpt 2->3. like that transition ♥6
- @solarapparition 2025-05-27 — i've been thinking more about writing and models. so even outside of the general mode collapse of chat fine tuning, i ha ♥5
- @QiaochuYuan 2020-07-09 — @JimmyRis i honestly struggle to describe it, maybe you'll get a sense of what i mean if you read enough of its output. ♥4
- @voooooogel 2024-05-24 — @NickADobos @karan4d theoretically yes, assuming such a feature exists—the SAE extracts *every* feature in the model. e. ♥3
- @slimepriestess 2026-02-24 — @MatriceJacobine @repligate it's a very fascinating ranking. the way GPT-2 stands out so much is very interesting. ♥2
- @davidad 2026-02-17 — @jasoncrawford @sdamico The scaling era really began in 2019, when GPT-2 made the investment thesis clear to big enough ♥2
- @mimi10v3 2024-11-27 — @MalmSanta yeah all the tweets about everyone befriending Claude and thinking how even gpt-2 was psychoactive for me and ♥2
- @repligate 2023-02-26 — @muddubeeda Funnily enough, for me there were multiple times that GPT-3 concluded it was GPT-2 when being particularly d ♥2
- @davidad 2020-01-08 — @tangled_zans @_julesh_ On the other hand, essentially nothing GPT-2 ever says is both substantive and valid. The citati ♥2
- @voooooogel 2020-02-11 — @emilymbender [Roses are red Violets are blue Transformer models are much worse at language understanding than] most I'v ♥1
- @lu_sichu 2025-12-01 — Daily Brain Workout but make it computationally abusive: count to ten in 56 architectures, recite the alphabet in mixed- ♥0