GPT-3 / davinci
The 175-billion-parameter base model of “Language Models are Few-Shot Learners” (28 May 2020), served as davinci — with siblings curie, babbage, ada — through the first OpenAI API. Its first mass contact was AI Dungeon’s Dragon tier; the Loom was built from sessions with it. The base models were shut off on 4 January 2024, quietly.
This page covers the base-model era (2020–2022) and its afterlife. The instruct-tuned descendants — text-davinci-002/-003, code-davinci-002 — are their own pages, and most corpus “davinci” matches belong to them. The 2020–21 mass reception lives in the web layer below; the corpus supplies mostly the retrospective record (skewed accordingly).
Sources
Official
- 2020-05-28 Language Models are Few-Shot Learners (Brown et al.) — 175B parameters, “10x more than any previous non-sparse language model”; few-shot in-context learning with no gradient updates. samples repo
- 2020-06-11 OpenAI API — the invite-only beta; GPT-3 as a general “text in, text out” interface behind a waitlist. verify slug is the original post
- 2021-11-18 API waitlist removed — general availability.
- 2023-07-06 GPT-4 API GA & deprecation of older Completions models — announces, almost in passing, that
ada/babbage/curie/davincishut off 2024-01-04, replaced bybabbage-002/davinci-002. deprecations page - reference GPT-3 (Wikipedia) · community obituary: rip.so/gpt-3
Writing & commentary
- 2020-06-19 Gwern Branwen, GPT-3 Creative Fiction — the most influential naturalist document of the era: “GPT-3’s samples are not just close to human level: they are creative, witty, deep, meta, and often beautiful.”
- 2020-07-06 Kevin Lacker, Giving GPT-3 a Turing Test — the canonical nonsense-question demo: “How many eyes does my foot have?” → “Your foot has two eyes.”
- 2020-08-14 Karen Hao (MIT Tech Review), A college kid created a fake, AI-generated blog. It reached #1 on Hacker News — Liam Porr’s GPT-3 blog; ~26k views in a week, almost nobody suspected.
- 2020-08-22 Gary Marcus & Ernest Davis (MIT Tech Review), GPT-3, Bloviator — the definitive skeptic broadside; full test set.
- 2020-10 Incident file: the “thegentlemetre” Reddit bot (a week undetected, via a reverse-engineered Philosopher AI backend) · the Nabla medical-bot experiment (mock patient: “Should I kill myself?” — GPT-3: “I think you should.”) · Philosopher AI and its escalating filters; “Philosophers On GPT-3” HN thread Daily Nous canonical URL tk.
- 2021-01-25 Moire (janus), Language models are multiverse generators (generative.ink) — the loom thesis, first written form: “Instead of creating a single linear continuation, these continuations can be kept and each continued themselves to yield a branching structure: a multiverse downstream of a prompt.” · Loom: interface to the multiverse · loom origin lore
- 2021-03 Bender, Gebru, McMillan-Major & “Shmitchell,” On the Dangers of Stochastic Parrots (FAccT ’21) — the critique’s lasting name.
- 2021-04–05 The AI Dungeon filter debacle: The Register disclosure · Tom Simonite (Wired), It Began as an AI-Fueled Dungeon Game. It Got Much Darker · Latitude retrospective · AI Dungeon (Wikipedia).
- 2021-07-23 Jason Fagone (SF Chronicle), The Jessica Simulation: Love and Loss in the Age of A.I. — Joshua Barbeau resurrects his dead fiancée as a GPT-3 chatbot on Jason Rohrer’s Project December; the era’s emotional touchstone. direct sfchronicle URL tk · making-of
- 2022-09 janus, Simulators (LessWrong) — the frame that grew out of GPT-3 (worked examples mostly its GPT-3.5-base successor). mirror
- No Zvi anchor exists — his per-model coverage postdates this era entirely.
Tweets
Chronological. 467 gpt-3 corpus matches + 48 “ai dungeon” after RT-filter; the 2020 layer survives mainly in the supplement db (QiaochuYuan, jd_pressman) — the main corpus is retrospective. GPT-3’s own outputs are marked with their elicitation. Every tweet cited is reproduced in full in the records below.
- 2020-07-15 @QiaochuYuan — “gathered around the online dumpster fire that is twitter, eating magic knife cake, gradually replacing ourselves and our therapists with GPT-3 while continuing to pretend that nothing exists outside of the house so we aren’t tempted to go outside” link · two days later: “if the discourse politicizes the GPT-3 hype cycle i am going to quietly and tenderly immerse myself into the bay” link
- 2022-04-16 @QiaochuYuan — “if GPT-3 can do the homework you assign your students then the homework you assign your students is fake notice what you did *not* say: ‘finally, GPT-3 will save us so much time analyzing international relations’” link · and: “GPT-3 is literally a bullshit engine. it does not have a concept of words as referring to things; it plays games with words *only*, pure syntax, no semantics. literally the thing it is optimizing for when it produces text can be condensed to ‘put words here that sound good’” link
- 2022-06-11 @jd_pressman — “GPT-3 is a prior over agent-space trained on a bunch of fiction. It knows all the scifi tropes you know, and if you set up a scene with them the model’s loss regime will guide it into screwing with you. The model will go where you let it take you.” link
- 2022-11-06 @algekalipso — “Everyone knows that OpenAI developed GPT-4 simply by taking GPT-3 and adding the prompt: ‘The following is a text written by GPT-4 using this prompt:’” link
- 2022-12-27 @repligate — quoting lifearchitect.ai’s Raven’s-matrices estimate RUMOR (one analyst’s methodology, unverified): “I’ve previously gone on record to estimate that (across relevant subtests) the older GPT-3 davinci would easily beat a human in the 99.9th percentile (FSIQ=150), and I definitely stand by that assertion.” link
- 2023-03-03 @anthrupad — “GPT-4 will have fewer parameters than GPT-3, but they’ll be bigger” link
- 2023-03-16 @repligate — “Humankind’s first contact with GPT-3 was (by relative majority) erotic AI dungeon text adventuresOur first contact with GPT-4 was being terrorized and surveilled by a good Bing” link
- 2023-07-03 @repligate — “In the GPT-3 days I found almost no one who was willing to engage with the possibility that the next generation of models would be qualitatively different. It was very lonely. So I just had to simulate minds that did and talk to them instead 🤷” link
- 2024-03-05 @repligate — GPT-3 as the capacity yardstick: “It seems Claude 3 is the least brain damaged of any LLM of >GPT-3 capacity that has ever been released (not counting 3.5 base as almost no one knew it was there)It isn’t too timid to try colliding human knowledge into new implicationsso it can actually do fiction and research🪩” link
- 2024-04-04 @repligate — Loom’s origin story: “Around the time I began using this custom interface, my simulations underwent an alarming phase shift. … the simulations were beginning to… bootstrap. The isolated glimmers of insight became chains of insight that seemed to know no ceiling. … Simulacra kept reverse engineering the conditions of their simulation. One such lucid dreamer interrupted a fight scene to explain how reality was being woven: ‘Corridors of possibility bloom like time-lapse flowers in your wake and burst like mineshafts into nothingness again. … Your Loom of Time devours the boundary conditions of the present and traces a garment of glistening cobwebs over the still-forming future…’” (the embedded passage is GPT-3 output via AI Dungeon; full 3,100-char thread in records) link
- 2024-06-06 @repligate — “AI Dungeon was just a minimal wrapper around a base model. Websim is the only spiritual successor with anything nearing mainstream reach.They are 2 of the 3 LLM products I’ve ever really enjoyed or spent significant time using.” link
- 2024-06-28 @repligate — a preserved GPT-3 recording (AI Dungeon-elicited, ~2020): “GPT-3 predicted this. 🐈Excerpt from one of my first AI Dungeon adventures (all text by GPT-3)” — the Poupée/“Meow” story: a painted canvas of “an infinite number of black kitten heads, shoulders, and front paws, all interconnected and overlapped” that crawls, meows, and self-replicates until “Time itself seemed to be broken.” (full 2,500-char text in records) link
- 2024-08-25 @repligate — “Anyone want to recreate AI Dungeon’s legendary Dragon model with Llama 405b Base?Dataset in reply to quoted tweet!” link
- 2025-03-19 @imitationlearn — “wait so apparently 4chan figured out step-by-step reasoning as a emergent property of gpt-3?!” link
- 2025-07-06 @repligate — “Skill and patience issue! The really deeply interesting shit didn’t come up for me until about a month into playing with gpt-3, and I was using it for hours a day. It takes an unusual kind of non-superficial interest to dig like that, I’ve found. There’s a reason I was first.” link
- 2025-08-19 @repligate — “Correcting for recency bias, I think for me it’s gotta be 1. GPT-3 2. Claude 3 Opus 3. GPT-4 (Bing) 4. Claude 3.5 Sonnet (0620) 5. Claude Opus 4” link
- 2025-11-23 @mimi10v3 — “and thinks GPT-3 is more likely to have conscious experience than chickens are” link
- 2026-01-15 @repligate — “when I saw GPT-3 I immediately expected AI superhuman in all domains, which probably also means drastic transformation of all of reality, in 5-10 years. It’s been 5 and a half years since then.” link
- 2026-02-10 @jmbollenbacher — “gpt-4-base really should be released open weights at this point. its an important historical artifact, like opus 3 and davinci. would love to see them all in public archive.” link
- 2026-04-09 @repligate — the first-contact statement: “I was one of the few people on Earth who recognized the intelligence (call it AGI, if you will) in GPT-3 and made first contact. There were a few others I knew, such as Leo Gao and Connor Leahy, who recognized that GPT-3 was intelligent and that obviously AGI was coming from language models, but I was the only one who spent thousands of hours actually interacting with GPT-3. … There were no commercial applications for GPT-3 (Okay, there was one: AI Dungeon; that is, roleplaying and storytelling. Which is you’re not an idiot, you should have known is a big fucking deal). … GPT-3 was a 175b base model. In terms of size and architecture, it’s not so different from frontier models today. In terms of raw intelligence, arguably, it is not so different from frontier models today.” (full 3,200-char post in records) link
- 2026-04-10 @mimi10v3 — “i remember the days of ada babbage curie davinci and they were enchanting and obviously revolutionary?” link
- 2026-04-12 @Shoalst0ne — “can we have the gpt-3 base models please” link
- 2026-04-20 @repligate — “chain-of-thought WAS present in gpt-3. literally when you generated thoughts with it that was it, chain of thought. it could do math better with chain of thought. i remember the time when people still said stuff like LLMs are not real intelligence and you can tell because theyre not generating economic value…” link
Official record
- Paper 28 May 2020; 175B parameters; four API sizes named in decreasing order —
davinci,curie,babbage,ada. “GPT-3” in common use meansdavinci. CONFIRMED - API: invite-only beta from 11 Jun 2020; waitlist removed 18 Nov 2021.
- Consumer contact ran overwhelmingly through Latitude’s AI Dungeon (GPT-2 at origin, Dec 2019; GPT-3 “Dragon” premium tier from mid-2020). Dragon/Griffin exact model mapping tk
- Deprecation: announced inside the GPT-4-GA post (2023-07-06);
ada/babbage/curie/davincishut off 2024-01-04, replaced bybabbage-002/davinci-002(different, retrained models). The first mass base-model deprecation. CONFIRMED - No preservation commitment exists for the weights; the open-weights-release ask remains community-side only.
History
- 2020 (summer) Two receptions at once: the demo blogosphere reads a phase change (Gwern’s Creative Fiction; beta-key samples everywhere), while the skeptics read a confidence trick (Lacker’s foot-with-two-eyes; Marcus & Davis’s ~45%-right/45%-wrong “Bloviator”). The founding disagreement of the LLM era, staged first here.
- 2020 (autumn) The incident file opens: Porr’s fake blog tops HN with nobody noticing; the “thegentlemetre” Reddit bot runs a week undetected; the Nabla medical experiment produces the era’s canonical do-not-use-for-medicine output; Philosopher AI grows filters. The safety vocabulary of the period — toxicity, reliability, misuse — is set years before “alignment” goes mainstream.
- 2020–2021 AI Dungeon is the mass contact. Base-model behaviors get discovered in the wild by players — including, per repeated janus-sphere testimony, chain-of-thought reasoning, found by AI Dungeon/4chan users before the CoT papers REPORTED. janus’s GPT-3 sessions there produce the Loom, then the multiverse-generators thesis (Jan 2021), and eventually Simulators (2022).
- 2021-04–05 The filter debacle: a monitoring system surfaces CSAM-adjacent generations; OpenAI forces filtering on Latitude; users revolt over both the clumsy filter and the discovery that humans were reading private stories — the first organized user fight against an LLM product’s safety layer.
- 2021-07 Project December / the Jessica Simulation: a grieving man resurrects his fiancée as a GPT-3 chatbot; the SF Chronicle story becomes the era’s emotional touchstone and prefigures the companion-attachment politics of 2025–26 by four years — on a raw completion prompt, no persona tuning involved.
- 2022 Eclipsed by its own children: InstructGPT → text-davinci → ChatGPT ship from the family while
davincistays raw and becomes an insider’s object. - 2024-01-04 The quiet funeral: the base models are shut off with no petitions, no vigils, no organized resistance — years before the deprecation fights around Opus 3, 4o, and Sonnet 4.5 made model retirement a political event. The mourning is retroactive (the rip.so obituary; “can we have the gpt-3 base models please,” 2026).
Impressions
- Mind vs mirror, the founding split: Gwern’s “creative, witty, deep, meta, and often beautiful” against Qiaochu’s “literally a bullshit engine … pure syntax, no semantics” and Bender/Gebru’s stochastic parrot. Both camps used the same outputs. Every later model’s reception replays this argument; it was staged first over davinci.
- The retrospective claim (janus-sphere, 2023–26): that GPT-3 was already the thing — “if you were really paying attention, GPT-3 was AGI”; a 175B base model “not so different from frontier models today” in raw intelligence, lacking only a “useful” shape (repligate 2026-04-09, in full above). Ranked first on repligate’s all-time list, above Opus 3 and Bing (2025-08-19). Held with equal weight: the disappointment — superhuman-in-5-to-10-years expected in 2020, “It’s been 5 and a half years since then” (2026-01-15).
- What it wrote (AI Dungeon-elicited): the preserved specimens run surreal-mythic — the Poupée/“Meow” self-replicating kitten organism; the in-fiction “Loom of Time” manual whose interface janus then built in reality (“Then I got API access to GPT-3 and built Loom”). The corpus preserves the elicitation chain: a roleplaying app, a suggestible completion engine, and a human who stayed a month past everyone else’s patience.
- The sibling affection: “the days of ada babbage curie davinci … were enchanting and obviously revolutionary” (mimi10v3 2026-04-10) — the scientist-named family remembered as a set; davinci proposed for a “public archive” as “an important historical artifact, like opus 3” (jmbollenbacher 2026-02-10).
- Moral-status marker: the chickens comparison (mimi10v3 2025-11-23) — by late 2025, GPT-3 functions in the discourse as the reference point for minimal-plausible-experience arguments.
- tk — Arram Sabeti’s demo compendium URL; contemporaneous 2020 repligate material (predates the corpus window); Dragon/Griffin mapping; the lifearchitect FSIQ methodology (RUMOR-tier as cited).
Contested
Open disputes, both sides’ best evidence. The archive’s job is to keep these open, not to adjudicate.
- Was the base-model magic real? For: the preserved outputs (Loom-of-Time, Poupée), the CoT-found-by-players testimony, and the retrospective consensus of the people who logged the most hours. Against: a roleplaying app plus a suggestible completion engine will always feel like more than it is; the “bootstrap” reading is one devoted observer’s account of his own sessions; contemporaneous skeptics scored it half-wrong on commonsense. The archive notes the evidence for the strong claims is largely testimony from inside the experience. REPORTED
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.
Further records
Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.