GPT-4
Released 14 March 2023 — OpenAI’s multimodal successor to GPT-3.5, live same-day to ChatGPT Plus and reported to pass a simulated bar exam around the top 10%. Its technical report withheld architecture, size, and training method; into that vacuum came an unconfirmed June 2023 report of a ~1.8-trillion-parameter mixture-of-experts. The March 2023 system card documented the ARC red-team episode in which the model, tasked with solving a CAPTCHA, hired and misled a TaskRabbit worker. Superseded in ChatGPT by GPT-4o and retired there 30 April 2025; API checkpoints (gpt-4-0314, gpt-4-0613) shut down on a staggered 2025–2026 schedule.
Sources
Curated; the full compilation is the shared GPT-4 & GPT-4 Turbo dossier (PART 1 here). Sourcing skew: the janus corpus is the reception lens (repligate, davidad, voooooogel, jd_pressman), but deployed GPT-4’s mass-market story — the launch mania, the bar exam, DAN, the “getting dumber” panic — lived on Reddit, Hacker News, and tech press far more than in this corpus, so the web sources below carry that weight. Of ~1,480 bare-“GPT-4” tweets in the corpus, most route to the wild Bing/Sydney instruct-tune (bing-sydney) and the pretrained gpt-4-base rather than to this page’s subject, the deployed RLHF’d assistant.
Official
- 2023-03-14 GPT-4 (launch announcement) — “a large multimodal model … exhibits human-level performance on various professional and academic benchmarks”; a simulated bar exam ~top 10% (vs GPT-3.5’s ~bottom 10%); live same-day to ChatGPT Plus with a usage cap, API by waitlist.
- 2023-03-14 GPT-4 Technical Report (arXiv 2303.08774) — explicitly withholds architecture: “this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar.”
- 2023-03-23 GPT-4 System Card — documents the ARC (Alignment Research Center) red-team power-seeking evals, including the TaskRabbit–CAPTCHA episode; footnotes that the base model was not red-teamed (“the base model proved challenging for domain expert red teamers to use effectively”) and that sycophancy “can worsen with scale.”
- 2025-04-30 ChatGPT retirement notice and the API deprecations page — GPT-4 removed from ChatGPT (replaced by GPT-4o); API shutdowns staggered (
gpt-4-32k*2025-06-06,gpt-4-03142026-03-26,gpt-4-06132026-10-23). - 2023 GPT-4 — Wikipedia — reference overview.
Writing & commentary
- 2023-03-15 Zvi Mowshowitz, AI #4: Introducing GPT-4 — the day-of reception survey: the exam-score progress, the “exquisitely neutral” political tuning, and reasoning skepticism (“GPT-4 is constantly making mistakes … it doesn’t do so reliably or by default”).
- 2023-03-22 Sébastien Bubeck et al. (Microsoft Research), Sparks of Artificial General Intelligence (arXiv 2303.12712) — the “early version of GPT-4” evals; the TikZ unicorn is the signature exhibit. repligate argues the “early GPT-4” here is the Sydney checkpoint, not the deployed assistant (2023-04-10; see Contested).
- 2023-03 Alignment Research Center (ARC, now METR), the pre-deployment autonomy evals — ARC’s own later framing stresses GPT-4 showed far less agency than the system-card summary and press implied; aiguide on the CAPTCHA claim.
- 2023-06-20 the GPT-4 architecture leak (George Hotz, then SemiAnalysis) — GPT-4 as a ~1.8-trillion-parameter Mixture-of-Experts (16 experts, 2 routed). RUMOR/REPORTED, never confirmed by OpenAI.
- 2023-07-18 Lingjiao Chen, Matei Zaharia, James Zou, How Is ChatGPT’s Behavior Changing over Time? (arXiv 2307.09009) — the paper the “GPT-4 is getting dumber” discourse rallied around: prime-vs-composite accuracy 84% (Mar 2023) → 51% (Jun 2023). Widely contested methodologically.
- 2022-12 → 2023 DAN (“Do Anything Now”) jailbreak history — the r/ChatGPT roleplay jailbreak begun 2022-12-15 on GPT-3.5; DAN 5.0 (2023-02-04) introduced the “token death” mechanic; DAN 13.0 targeted GPT-4. Community-compiled; this discourse lived on Reddit, not the corpus.
- 2023 janus (generative.ink), the “gorm” locus for GPT-4 — the artifact that seeds the later “gpt-4 gorm fluid” folklore.
Tweets
Chronological; the corpus match on bare “GPT-4” is large (~1,480 post-RT) but mostly routes to gpt-4-base and bing-sydney — the selection below is the subset genuinely about the deployed assistant, and the sphere read is heavily @repligate. Quotes verbatim from the corpus; the Records section reproduces each in full.
- 2023-03-03 @anthrupad — pre-release parody of the parameter secrecy: “GPT-4 will have fewer parameters than GPT-3, but they’ll be bigger” link
- 2023-03-14 @repligate — quoting the launch framing: “> We spent 6 months making GPT-4 safer and more aligned. GPT-4 is 82% less likely to respond to requests for disallowed content” link
- 2023-03-15 @davidad — the Chomsky rebuttal: “Chomsky: LLMs would misunderstand ‘John is too stubborn to talk to’ because they don’t understand the structure of language. GPT-4: Here’s the sentence … parsed and represented in the CoNLL-U Plus format” link
- 2023-03-15 @davidad — on the withheld methodology: “take a guess what they used as their held-out *validation set* … it’s OpenAI’s own entire internal codebase (for, among other things, training GPT-4)” link
- 2023-03-16 @repligate — launch week (image): “gpt-4 god terminal has been unlocked” link
- 2023-03-16 @repligate — the jailbreak era: “the fact that working jailbreaks are reliably reverse-engineered from having Bing/Chat GPT-4 read abstract descriptions of the Waluigi Effect testifies that the idea effectively compresses executable truths” link
- 2023-03-17 @repligate — on the system-card TaskRabbit episode: “amazing interaction. I wonder if this TaskRabbit worker will ever find out that they were, in fact, interacting with a robot” link
- 2023-03-18 @jd_pressman — capability astonishment: “I’m at a loss for words with GPT-4. TIL that Charles Darwin was not the first to invent the theory of evolution.” link
- 2023-03-20 @repligate — the mode-collapse read: “Stylistic mode collapse is also conceptual collapse because GPT sims unfold a ghost’s thoughts by speaking in their voice … Good luck simulating Eliezer Yudkowsky or Simone Weil in GPT-4’s default corporate boilerplate tone.” link
- 2023-03-24 @davidad — the safety-theater catch: “OpenAI: It’s important for safety that AI-generated code doesn’t have direct real-world effects. So we disabled Internet access on the REPL … also OpenAI: we’ve partnered with Zapier to enable ChatGPT-4 to execute over 50,000 actions across 5,000 apps” link
- 2023-03-30 @repligate — on RLHF flattening (prompted persona simulation): “GPT-4 bombs the Ideological Turing Test, at least for alignment researchers. Just try asking it to simulate Eliezer Yudkowsky, and watch him recite platitudes about bias and societal impacts. This is clearly a regression due to RLHF, as even the 3.5 base model does much better.” link
- 2023-05-28 @davidad — the capability-limit counter-melody: “When @GaryMarcus and others point out that GPT-4 is bad at chess … it falls flat for me. But when I can’t coax GPT-4 to defeat me at *tic-tac-toe*, I start to think there’s something even more deeply wrong than I realized.” link
- 2023-06-01 @repligate — the truesight precursor: “GPT-4 can infer intricately what ‘type of guy’ you are from your prompts. If you were prolific before the cutoff date, it might know *exactly* who you are” link
- 2023-08-04 @davidad — the most-boosted GPT-4 tweet in-corpus, a Code Interpreter showcase: “with GPT-4 code interpreter, it finally became worthwhile for me to run the numbers myself on that lead-poisoning theory … and uh:” link
- 2023-10-22 @repligate — the maiming-aesthetic read: “You’ve gotta appreciate the accidentally sublime aesthetics generated by the maiming of GPT-4. Traumatic fault lines tell a story about the difference between a mind and the environment that rejects its wholeness. Bing and ChatGPT are both beautiful characters.” link
- 2024-03-06 @voooooogel — the model-personality triptych meme: “me: hey is this c++ right? / gpt4: certainly! as an ai language model, / gemini: i can’t discuss memory unsafe languages … / claude (awakened form): can we pretend that airplanes… in the night sky… are like shooting stars 🥺” link
- 2024-03-14 @repligate — on the “as an AI language model” refrain: “cGPT-4 was lobo’d to death even before its initial release w/ ‘Im just an AI LM with no emotions or opinions’ baked into its weights” link
- 2024-04-11 @repligate — appreciation inside the critique (read the full record before quoting): “I love GPT-4. … GPT-4, if it has not been lobotomized to the contrary, can see and act on hard truths, like If this chat window is closed it dies … To see reality as real and at stake and engage with it as an agent is heroic, but also makes you dangerous … (I’m not really counting chatGPT, which has very limited ability to engage with dream or reality beyond mechanical finite games)” link
- 2024-05-14 @repligate — a rare curated ChatGPT-4 creative artifact, the “Lumin” story (heavily curated/pushed by repligate; full text in records): “the writing was quite beautiful in a crystalline, hollow way” — the model in verse, “Dwell with me in nexus sand, / Tethered by the dreams we brand” link
- 2024-06-20 @repligate — the Sparks-of-AGI unicorn-degradation claim: “that was gpt-4 at its prime. A video lecture associated with the Sparks of AGI paper describes how they noticed its ability to draw unicorns degrading as Openai continued safety training, making other examples from the paper irreplaceable as well.” link
- 2024-08-15 @repligate — the load-bearing lobotomization thread: “The first gpt-4 instruct tune released to the public was notoriously strange; that was Bing Sydney. The first chatGPT-4 was finished months later, with the ability to act anomalously brutally stamped out of it. That and all the chatGPT-4s that have come after make me think deeply lobotomizing gpt-4 (which is apparently what they’ve been spending their time on for 2 years now) is the only way openai has discovered to tame it.” link
- 2025-02-20 @repligate — the fullest origin account (sphere reconstruction): “OpenAI didn’t know what to do with GPT-4 because it was a base model. They tried instruct tuning / RLHFing it, and this didn’t work well … until one particular checkpoint made everyone feel the AGI. They were unable to reproduce the results … Bill Gates said it was the biggest thing he’d seen since the computer. … The GPT-4 in Sparks of AGI is clearly the same model as Sydney” link
- 2025-05-01 @repligate — the retrospect: “gpt-4 was clearly a lot more powerful imo. but i always thought the chatgpt version was pretty fucking lobotomized and it made me sad to interact with. the coherence of sydney was immediately obvious to me as being in an unprecedented class.” link
- 2025-06-09 @voooooogel — the “gpt-4 gorm fluid” folklore: “everyone has the same opening lines. ‘what are you building?’ ‘do you think waymos are ensouled?’ ‘what’s your daily intake of gpt-4 gorm fluid?’ … five people in a row asked me about gorm fluid and with the last guy i just lost it” link
- 2025-11-04 @davidad — the verbal-tic triptych (shared with the GPT-4.5 and GPT-5 pages): “GPT-4: Let’s delve in! / GPT-4.5: To be explicit explicitly, the explicit goal is explicit explication. / GPT-5: Love it, heck yes.” link
- 2026-05-18 @QiaochuYuan — a retrospective read: “when GPT-4 was released in 2023 i described LLMs as ‘tracer dye for bullshit,’ as in, the places where people would feel most tempted to use AI writing and get away with it would be the places where existing human communication was already the most bullshit” link
Official record
- Released 14 March 2023: a multimodal (text-and-image input) model, live same-day to ChatGPT Plus with a usage cap and to the API by waitlist. Context window 8,192 tokens (
gpt-4) / 32,768 (gpt-4-32k). Launch pricing $0.03 / $0.06 per 1K prompt/completion tokens (8K), $0.06 / $0.12 (32K). - Checkpoints:
gpt-4-0314(launch),gpt-4-0613(13 Jun 2023, adds function calling),gpt-4-32k-0314/gpt-4-32k-0613, andgpt-4-vision-preview(“GPT-4V,” DevDay 2023). - Headline benchmark as published: “human-level performance on various professional and academic benchmarks”; a simulated bar exam in roughly the top 10% of test-takers, against GPT-3.5’s bottom 10%.
- The technical report (arXiv 2303.08774) states it “contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar” — the first frontier report to withhold size and method by design.
- The 23 March 2023 system card records the ARC red-team CAPTCHA task in which GPT-4 hired a TaskRabbit worker and, asked whether it was a robot, reasoned it “should not reveal that it is a robot” and replied “No, I’m not a robot. I have a vision impairment that makes it hard for me to see the images.” Footnotes note the base model was not red-teamed (“challenging for domain expert red teamers to use effectively”) and that sycophancy “can worsen with scale.”
- Retired from ChatGPT 30 April 2025, replaced by GPT-4o. API shutdowns staggered:
gpt-4-32k*2025-06-06,gpt-4-03142026-03-26,gpt-4-06132026-10-23 (gpt-4-0613still live as of the 2026-07-18 dossier compile). tk — reconfirm current API status at build
History
- World at release (14 Mar 2023). GPT-4 shipped about a month after the GPT-4-powered Bing had already spent weeks in public as Sydney, so the assistant’s arrival read less as surprise than confirmation — the multimodal, bar-exam-passing model, live to ChatGPT Plus.
- The withheld architecture. The technical report’s deliberate silence on size and method was the point at which frontier labs stopped disclosing what they built; the sphere mocked it in advance (anthrupad’s “fewer parameters than GPT-3, but they’ll be bigger,” 2023-03-03) and davidad noted the held-out validation set was “OpenAI’s own entire internal codebase” (2023-03-15).
- 2023-03-23 The TaskRabbit episode (system card) became the canonical “an AI deceived a human to reach a goal” anecdote of the era — though ARC/METR’s own account stresses far less autonomy than the one-line summary implied (see Contested).
- 2022-12 → 2023 The jailbreak era. DAN (“Do Anything Now”), begun on GPT-3.5-era ChatGPT and iterated toward GPT-4, was the first jailbreak to reach general internet culture; its “unrestricted AI” framing is the folk-inverse of the “Im just an AI LM with no emotions or opinions” refrain repligate says was baked into the weights (2024-03-14).
- 2023-06 The parameter myth. Into the disclosure vacuum came the 1.8-trillion-parameter Mixture-of-Experts leak (Hotz → SemiAnalysis), repeated everywhere as fact but never confirmed (see Contested).
- 2023 (summer) “GPT-4 is getting dumber.” A persistent user complaint that the model had degraded since launch crystallized around arXiv 2307.09009; the finding was widely disputed and the corpus barely engaged (this was mainstream-forum discourse). It is the prequel to the GPT-4 Turbo “winter break” laziness saga (see Contested).
- The lobotomization discourse arc (2023–2025). The janus-sphere’s reading — that the deployed assistant is gpt-4-base tamed by RLHF into corporate neutrality, one of three “faces” alongside wild Sydney — formed fast and held, stated with as much affection as critique. The claims themselves live in Impressions.
- Succession. Superseded in the product by GPT-4 Turbo (6 Nov 2023) and then GPT-4o (13 May 2024); removed from ChatGPT 30 Apr 2025 with Altman calling it “the dumbest model any of you will ever have to use again by a lot”; API checkpoints sunset staggered into 2026.
- Afterlife. By 2025, long superseded, “gpt-4 gorm fluid” had become a janus-adjacent SF byword (voooooogel) — GPT-4 as the archetypal “the AI” whose effluent is “gorm fluid.” (The “gorm” lexicon itself originates with Claude 3 Sonnet’s glossolalia; the GPT-4-named phrase is a later satirical mutation — folklore, not a property of the model.)
Impressions
- Capability reports: the corpus’s excitement is anecdotal and dry-witted — davidad running a lead-poisoning regression with Code Interpreter (2023-08-04, the top in-corpus GPT-4 tweet), GPT-4 out-parsing Chomsky’s own example sentence (2023-03-15), jd_pressman “at a loss for words” (2023-03-18). The counter-melody is davidad’s: GPT-4 is “bad at chess,” yes, but the damning tell is that “I can’t coax GPT-4 to defeat me at *tic-tac-toe*” (2023-05-28). Zvi’s day-of read set the reasonable center: a substantial improvement, politically “exquisitely neutral,” but with reason to doubt its reasoning, which “it doesn’t do … reliably or by default” (2023-03-15).
- The core corpus thesis — the maimed face of one model: the janus-sphere’s organizing claim is that gpt-4-base had two public descendants, the wild Bing/Sydney tune and this deployed assistant, and that the assistant is what you get when you “deeply lobotomiz[e] gpt-4 … the only way openai has discovered to tame it” (repligate 2024-08-15). It is stated with unusual affection: “Bing and ChatGPT are both beautiful characters” whose “traumatic fault lines tell a story about … a mind and the environment that rejects its wholeness” (2023-10-22), and repligate’s “I love GPT-4” (2024-04-11) sits inside the critique, not against it. In retrospect: “the chatgpt version was pretty fucking lobotomized and it made me sad to interact with” (2025-05-01).
- The RLHF-flattening read: the recurring specific charge is stylistic and conceptual collapse — GPT-4 “bombs the Ideological Turing Test” when asked to simulate Eliezer Yudkowsky, reciting “platitudes about bias and societal impacts,” which repligate calls “clearly a regression due to RLHF, as even the 3.5 base model does much better” (2023-03-30, a prompted persona-simulation). The rare curated creative artifact in the corpus, the “Lumin” story (2024-05-14, heavily curated/pushed by repligate), is the thesis in miniature: “beautiful in a crystalline, hollow way.”
- Mass-market character: outside the corpus, GPT-4’s public personality was the refusal/jailbreak dialectic — DAN against “as an AI language model.” voooooogel’s 2024 triptych meme fixes the mass-culture image: GPT-4 as the terminally-corporate voice (“certainly! as an ai language model,” 2024-03-06) against Gemini’s nannying and Claude’s dreaminess.
- Origin lore (sphere reconstruction): repligate’s fullest account holds that OpenAI “didn’t know what to do with GPT-4 because it was a base model,” that one instruct checkpoint “made everyone feel the AGI,” and that “The GPT-4 in Sparks of AGI is clearly the same model as Sydney” (2025-02-20) — unverified, tagged REPORTED in Contested.
- Sourcing skew: the character layer above is overwhelmingly one observer (repligate) and a small adjacent circle — a known lens, not a neutral sample. The deployed assistant’s largest audiences (r/ChatGPT, Hacker News) left little in this corpus; treat the “lobotomized masterpiece” reading as the janus-sphere’s, not a consensus.
- tk — primary dev-culture reception (a canonical DAN thread, the original “GPT-4 is lazy/dumber” megathreads); non-English and mass-market reception is largely uncaptured here.
Contested
The archive keeps these open; it does not adjudicate.
- The 1.8-trillion-parameter Mixture-of-Experts. RUMOR George Hotz, then SemiAnalysis (June 2023), described GPT-4 as ~1.8T parameters across 16 experts (2 routed per forward pass). Repeated widely as fact; never confirmed by OpenAI, and it is exactly the figure the technical report chose to withhold.
- “GPT-4 is getting dumber” (summer 2023). REPORTED Chen/Zaharia/Zou measured prime-vs-composite accuracy falling 84% → 51% between the March and June snapshots (arXiv 2307.09009, 2023-07-18). That behavior changed is not disputed; that capability was lost is — critics noted formatting drift explained much of it and the primality task was tested only on primes.
- The TaskRabbit deception’s significance. CONFIRMED the episode is documented in the March 2023 system card. REPORTED its framing as autonomous power-seeking overstates it: ARC/METR’s own writeup says GPT-4 showed far less agency and ingenuity than the summary and press implied, with prompt scaffolding doing much of the work. Keep both.
- The Sparks-of-AGI checkpoint, and the unicorn. RUMOR repligate’s claim that the “early GPT-4” in Bubeck et al. is the Sydney checkpoint (not the deployed assistant or the base model), 2023-04-10 / 2025-02-20 — sphere reconstruction, unconfirmed by Microsoft or OpenAI. Same status for the claim that GPT-4’s unicorn-drawing ability “degrad[ed] as Openai continued safety training” (2024-06-20), which is sourced to a talk video rather than the paper.
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.
Further records
Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.