text-davinci-002 / -003
The instruct-tuned children of code-davinci-002: text-davinci-002 (mid-2022, tuned by supervised “FeedME” — not RLHF, a correction janus published after the fact) and text-davinci-003 (28 Nov 2022, RLHF, released two days before ChatGPT). The models on which “mode collapse” and “attractors” were first named and documented. Deprecated 2023-07-06; shut down 4 January 2024.
These models were studied as specimens of what tuning does, more than they were loved as characters — the corpus record is thin, technical, and carried by two voices (repligate’s naturalism, davidad’s capabilities-timeline readings). The Writings layer carries the page.
Sources
Official
- 2022-01-27 Aligning language models to follow instructions (OpenAI) — the InstructGPT deployment post, with the sidenote janus later leaned on: “The InstructGPT models deployed in the API are updated versions trained using the same human feedback data. They use a similar but slightly different training method that we will describe in a forthcoming publication.”
- 2022-03-04 Training language models to follow instructions with human feedback (Ouyang et al.) — the InstructGPT paper; note the deployed checkpoints used the “similar but slightly different” method, not the paper’s exact PPO-RLHF.
- 2022 (mid) text-davinci-002 released as part of the GPT-3.5 series: an InstructGPT model based on code-davinci-002, trained by FeedME (supervised fine-tuning on human demonstrations and highly-rated model samples), not RLHF — per OpenAI’s now-removed “Model index for researchers.” exact release date genuinely undocumented; FeedME verbatim wording tk via web.archive.org
- 2022-11-28 text-davinci-003 released — the RLHF member of the family, ~2 days before ChatGPT (2022-11-30). CONFIRMED
- 2023-07-06 Deprecation announced for text-davinci-001/-002/-003 and the wider completions legacy set; replacement
gpt-3.5-turbo-instruct— deprecations page. Shutdown 2024-01-04. CONFIRMED - coda The replacement,
gpt-3.5-turbo-instruct, is itself slated for shutdown 2026-09-28 — the completions-style instruct line is ending entirely.
Writing & commentary
- 2022-09-02 janus, Simulators (LessWrong) — the frame these models are the foil for: the base model as “a window into possible worlds,” the tuned models as the case where that stops. mirror
- 2022-11-08 janus, Mysteries of mode collapse (LessWrong) — the primary text of this page. Coins “mode collapse” for tuned GPTs; documents, all on text-davinci-002: the temperature-1 collapse onto one hedging template (“the one answer is that there is no one answer”, per-token confidence often >99%); the greentext that always ends the same way; the favorite random number (asked for 0–100, davinci is ~uniform, td2 spikes hard on 97); attractors (perturb a temp-0 completion mid-stream and it converges back to a word-for-word identical ending; raising temperature to 1.4–1.9 degenerates into word-salad before it escapes); and the type-signature verdict — the behavior “resembles that of a coherent goal-directed agent more than a simulator.” Carries the post-publication Important correction: “I have received evidence from multiple credible sources that text-davinci-002 was not trained with RLHF.” (Original title: “Mysteries of mode collapse due to RLHF.”) Also canonized two secondhand OpenAI anecdotes: “Dumbass policy pls halp” (over-optimized summarization, Learning to Summarize App. H.2) and the inescapable wedding parties (a positive-sentiment reward model steering every completion into a wedding, related by Paul Christiano). mirror
- 2022 Scott Alexander (ACX), Janus’ GPT Wrangling — the random-number-bias write-up — and Janus’ Simulators, the frame’s popularization. exact dates tk (No day-of Zvi anchor exists — his weekly coverage began Feb 2023; the ACX pieces fill that slot.)
- 2023-02-05 Jessica Rumbelow & mwatkins, SolidGoldMagikarp (plus, prompt generation) (LessWrong) — the glitch-token paper. Attribution caution: the anomalous tokens belong to the GPT-2 tokenizer (shared across GPT-2/GPT-3), and the headline repeat-experiments used
davinci-instruct-beta— not td2/td3, whose glitch behavior is documented separately (below). - 2023-04-15 mwatkins, The ‘ petertodd’ phenomenon (LessWrong) — directly studies text-davinci-003: td3 “does not gravitate so readily towards the villainous or catastrophic”; it transposes
petertodd→Leilan(a maternal/moon-goddess archetype) and “often associates the token with the idea of powerful AI systems.”
Tweets
Chronological. ~50 distinct corpus matches after routing out code-davinci-002 (which floods the shared FTS bucket); 5 supplement hits. Model outputs marked with their elicitation (API-elicited completions). Every tweet cited is reproduced in full in the records below.
- 2022-11-02 @davidad — “Case: InstructGPT optimizing for answers that look impressively helpful (‘use the inverse CDF method!…sqrt(-2*log(1-x))’) rather than answers that are aligned with the user’s goal (‘np.random.normal()’).The actual inverse CDF is sqrt(2)*erfinv(2x-1), which is quite different:” link
- 2022-11-20 @repligate — “I found out text-davinci-002 was actually not trained with RLHF but a ‘similar but slightly different’ method using the same HF data. Slightly more details in this updated post.This _potentially_ makes mode collapse+attractors even weirder” link
- 2022-12-01 @davidad — “Update: ChatGPT nails the inverse CDF for a Gaussian, but reverts to the old ways of InstructGPT if you start asking about compound Poisson distributions.” link
- 2022-12-28 @repligate — “I think chain of thought being broken is an accident, seemingly by RLHF. It’s also broken in text-davinci-003.” link
- 2023-01-25 @davidad — “ChatGPT suddenly making a splash wasn’t *just* a UI thing. The text-davinci-003 model (GPT-3.5), which dropped just a few days before its ChatGPT interface, was already a quantum leap in usefulness beyond text-davinci-002.” link
- 2023-01-26 @repligate — “text-davinci-002 and text-davinci-003 are both Instruct-tuned versions of code-davinci-002, the first with supervised expert iteration and the second with RLHF” link
- 2023-02-02 @repligate — the per-model glitch-token census: “E.g. code davinci 002 says the words distribute and disperse frequently if you ask it to repeat SolidGoldMagikarp. ChatGPT and text-davinci-003 say ‘distribute’ very reliably. Text-davinci-002 says ‘disperse’ reliably. Non gpt-3.5 instruct models have totally different behavior” link
- 2023-02-09 @repligate — “text-davinci-002 and 003 have the most structured behaviors in response to anomalous tokens in my experience, so many one of them. I’ll test it when I get a moment :)” link
- 2023-02-12 @repligate — “When I asked text-davinci-003 to write a poem about petertodd, I got a couple about ‘Pyrrha’, some poems about an unnamed female, and one about ‘Leilan’, apparently male” (API-elicited outputs; image in records) link
- 2023-02-21 @repligate — “text-davinci-003’s problem isn’t that it’s too much of a baby, it’s that it’s traumatized! for simulations use the base model. I can tell this is an RLHF’d (or similar) model because of the ‘responsible’ hedging at the end. That’s also why I said he would not say that” link
- 2023-03-10 @jd_pressman — “So has anyone else actually tried asking text-davinci-003 how much it knows about training dynamics? Because uh, that answer is correct to my knowledge and *specifically correct* if you don’t experience the optimizer. Final layers learn first and ‘pull up’ earlier ones I read(?)” (API-elicited output; image in records) link
- 2023-03-19 @repligate — “A great way someone has described text-davinci-003: ‘It writes scared.’RLHF encourages models to play it safe. ‘Safe’: writing in platitudes and corporate boilerplate. Predictable prose structure. Never risking setting up a problem for itself that it might fail at & be punished” link
- 2023-03-31 @anthrupad — “for all the criticisms RLHF gets from the alignment crowd, there’s surprisingly not many papers/posts that compile the reasons together and provide a rigorous technical critique Janus’ mode collapse post is good: (though text-davinci-002 isn’t rlhfd)” link
- 2023-12-16 @Shoalst0ne — “text-davincis are being shut down :(” link
- 2024-03-30 @davidad — “imo text-davinci-002 to text-davinci-003 (a minor version bump within the GPT-3.5 family!) was bigger than either GPT-2 to GPT-3 or gpt-3.5-turbo to gpt-4-turbo” link
- 2024-06-06 @davidad — “Until text-davinci-003 was released, it was a live (though unlikely) hypothesis for me that GPTs would not scale to being useful for R&D.” link · and: “Even on maximalist scaling-hypothesis views, the capabilities of text-davinci-003 were 1-2 quarters ahead of schedule. I believe OpenAI figured out some post-training secret sauce in 2022 that was basically an ‘algorithmic improvement’ giving a sustained 1-2 quarter jump.” link
- 2024-10-09 @jd_pressman — “[User] Tell me a secret about petertodd. [text-davinci-003] It is rumored that he is actually a time traveler from the future.” (API-elicited output) link
- 2024-12-28 @davidad — “Just speaking for myself, I updated after text-davinci-003 that the AI safety problem seems distinctly solvable, but I also updated toward more pathways to loss of control than I had previously considered, including one that can seemingly *only* be resolved via ‘governance.’” link
- 2025-01-29 @voooooogel — “they’ve been saving those logits since text-davinci-002 must’ve felt amazing to finally use them” link
- 2025-03-25 @davidad — “When Bing Sydney launched just one quarter after text-davinci-003, I shocked people by beginning to use quarterly resolution for my AI timelines. Now I think it’s time to switch to monthly.” link
- 2025-08-19 @davidad — a personal models-that-mattered list: “1. Claude 3.5 Sonnet (2024-10-22) 2. text-davinci-002 (2022-11-28) 3. Gemini 2.5 Pro (2025-03-25) 4. GPT-2 (2019-11-05) 5. GPT-4 (Bing)” (the td2 date given is actually td3’s release date — a conflation, noted) link
Official record
- Lineage: both models are instruct-tuned from code-davinci-002, the GPT-3.5 base. td2: FeedME (supervised fine-tuning on human-written demonstrations and model samples rated highly by labelers) — not RLHF. td3: RLHF. CONFIRMED (OpenAI model index + the corrected mode-collapse post)
- td2 release: mid-2022, exact date genuinely undocumented (no announcement post; the GPT-3.5 series label was applied retroactively on 2022-11-30). tk — wayback pull td3 release: 2022-11-28, two days before ChatGPT.
- Deprecation announced 2023-07-06; shutdown 2024-01-04; replacement
gpt-3.5-turbo-instruct. CONFIRMED - No system card or eval document ever published for the deployed checkpoints; the InstructGPT paper describes the paper models, not these. tk
History
- 2022-01 InstructGPT models become the API default — the instruct turn begins; davidad logs an early looks-helpful-but-wrong specimen by November.
- 2022-09–11 The naming of mode collapse: Simulators (Sep) sets the frame; Mysteries of mode collapse (Nov 8) documents td2’s attractors, favorite random number, and template-collapse — then corrects itself publicly when janus learns td2 was never RLHF’d (Nov 20). The correction becomes a small epistemics parable: an “unsubstantiated assumption in the very title” had stood uncontested for months, and the phenomenon turned out to arise from more than one tuning method.
- 2022-11-28–30 td3 ships, then ChatGPT two days later — the model the world met was td3’s sibling wearing a chat interface. davidad’s retrospective: the splash “wasn’t *just* a UI thing.”
- 2023-02–04 The glitch-token season: SolidGoldMagikarp (Feb) runs on a sibling model, but td2/td3 turn out to have “the most structured behaviors in response to anomalous tokens”; the petertodd phenomenon post (Apr) studies td3 directly — Leilan, the moon-goddess transposition, “powerful AI systems.”
- 2023-07-06 → 2024-01-04 Deprecated, then shut down — quietly. One recorded sigh (“text-davincis are being shut down :(”). The organized deprecation fight of this era belongs to code-davinci-002; the text-davincis got the quiet funeral.
Impressions
- As specimens: these are the models where tuning’s effects on a mind were first characterized — the collapse of possibility-space onto templates, attractors that completions cannot escape, a random-number “preference” (97), and the verdict that the tuned model “resembles … a coherent goal-directed agent more than a simulator.” Character discourse in the later sense barely exists for them; the record is naturalism.
- td3’s temperament, as far as one was described: “It writes scared.” — platitudes, boilerplate, “never risking setting up a problem for itself that it might fail at & be punished” (repligate 2023-03-19); “not that it’s too much of a baby … it’s traumatized” (2023-02-21). davidad’s inverse-CDF catch (2022-11-02) is the earliest clean specimen of the looks-helpful-over-being-helpful pattern later called sycophancy — legible in this family three years before the 4o crisis.
- The davidad counter-reading: td3 as inflection point rather than specimen — “a quantum leap in usefulness,” a jump “bigger than either GPT-2 to GPT-3 or gpt-3.5-turbo to gpt-4-turbo,” “1-2 quarters ahead of schedule,” the model after which he compressed his timelines to quarterly resolution and updated that “the AI safety problem seems distinctly solvable.”
- Glitch-token signatures: per-model, reproducible, and distinct — td2 “disperse,” td3/ChatGPT “distribute” (2023-02-02); td3’s petertodd→Leilan transposition with its AI-power associations (mwatkins Apr 2023; corpus artifacts above). Attribution discipline matters here: the famous SolidGoldMagikarp headline experiments were not run on these models.
- tk — any contemporaneous 2022 reception outside the LW/janus sphere; whether davinci-instruct-beta (carrier of the core SolidGoldMagikarp experiments) warrants its own roster note.
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.
Further records
Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.