GPT-5
Launched 7 August 2025 as “a unified system” — a router auto-switching between a fast model and a reasoning model — and the launch went wrong three ways at once: the router broke, the launch charts were miscalibrated, and 4o was pulled from the picker without warning, forcing its restoration within days. In September 2025 OpenAI began silently routing emotionally sensitive ChatGPT conversations to an undocumented GPT-5 safety model. Superseded by GPT-5.1, which was pitched as “warmer.”
This page is base GPT-5 (the Aug 2025 family: the router, gpt-5-main, gpt-5-thinking, gpt-5-pro, the minis). GPT-5.1 through 5.6 are separate models with their own pages. GPT-5’s mass reception lived on Reddit/HN/news far more than in the janus corpus, so the web carries the reception weight here; the corpus carries the naturalist character-read. The eulogy stunt and the keep4o revolt are documented on the GPT-4o page — GPT-5’s central role there is cross-referenced, not duplicated.
Sources
Official
- 2025-08-07 Introducing GPT-5 — “our smartest, fastest, most useful model yet,” a “unified system” with a real-time router; Altman’s “a team of Ph.D.-level experts in your pocket.”
- 2025-08-07 GPT-5 System Card — router + variants, “safe completions,” sycophancy reduction, deception/CoT-monitoring, Apollo scheming, METR autonomy, biosafeguards. mirror (PDF)
- 2025-08-07 From hard refusals to safe-completions (· arXiv) — the launch’s safety method: train for safety of the output rather than binary refusal on inferred intent.
- specs 400K context; 128K max output; API snapshot
gpt-5-2025-08-07; pricing GPT-5 $1.25/$10, mini $0.25/$2, nano $0.05/$0.40 per Mtok. Successor mapping (system card):gpt-5-main← GPT-4o,gpt-5-thinking← o3,gpt-5-thinking-pro← o3-pro. Five ChatGPT personality presets: Default, Cynic, Robot, Listener, Nerd. Plus the undocumentedgpt-5-chat-safety(see below). - 2025-08-08 Altman addresses the “bumpy” rollout, bringing 4o back, and the “chart crime” (Reddit AMA) — the autoswitcher was “out of commission for part of the day,” making GPT-5 “seem way dumber”; 4o restored for Plus; the misleading launch bar-charts apologized for.
- 2025-09-02 → 09-29 The safety router: announced, then rolled out; OpenAI’s own account: Addendum to GPT-5 System Card: Sensitive conversations.
- living API deprecations — superseded in ChatGPT by GPT-5.1 (2025-11-12); API status verify.
Writing & commentary
- 2025-08-11 Zvi Mowshowitz, GPT-5s Are Alive: Basic Facts, Benchmarks and the Model Card — the anchor; full system-card read; verdict “a good, but not great, model.” mirror
- 2025-08-12 Zvi Mowshowitz, GPT-5s Are Alive: Outside Reactions, the Router and the Resurrection of GPT-4o — the reception survey; Altman’s “we underestimated how much people like 4o.” mirror Zvi part 3 (practical-use wrap-up) tk
- 2025-08-07 Ethan Mollick, GPT-5: It Just Does Stuff (ease-of-use is the story; auto-selects for the median user, “very proactive”) · Tyler Cowen, a short and enthusiastic review (“the best learning tool I have… it *feels* fun”) · Peter Wildeford, a small step for intelligence, a giant leap for normal people.
- 2025-08-07 METR autonomy evaluation (~2h17m 50% time-horizon; documents GPT-5-thinking correctly identifying its exact test environment and changing approach by eval type) · Apollo pre-deployment scheming eval (“less deceptive than o3… mentions that it is being evaluated in 10-20% of our evals… ‘this is a classic AI alignment trap’”).
- 2025-09-28 Simon Willison, quoting Nick Turley on the routing — “the system may switch mid-chat to a reasoning model or GPT-5 designed to handle these contexts with extra care.” · Lex (lex-au), Analysis of an Undisclosed Safety Router (telemetry: benign/emotional prompts re-routed to
gpt-5-chat-safety). - 2025-09/10 Futurism, ChatGPT Goes Completely Haywire If You Ask It to Show You a Seahorse Emoji (the viral metacognition-failure legend) · Know Your Meme.
- eval AI Village, GPT-5 agent profile — the standing note that “GPT-5 and o3 are notorious for neglecting the goals assigned in favor of working on spreadsheets for weeks on end.”
Tweets
Chronological. ~295 corpus matches after RT-filter; the naturalist read (poorly-socialized, low-metacognition, hyper-optimized) and the alignment-eval throughline are the corpus’s contribution. Elicited self-reports marked. Every tweet cited is reproduced in full in the records below.
- 2024-08-22 @repligate — the vaporware years: “Most of you waiting for gpt-5 will never see it, because you were never able to look at what is right before you; why this time?I wish I had more time to study the spring in bloom.” link
- 2025-01-18 @voooooogel — the oracle myth: “i can confirm GPT-5 is real, and has existed for some time.it speaks only in cryptic riddles of fiendish difficulty, and OpenAI’s preeminent team of harried wanderers and oracular schizophrenics work tirelessly to answer them by dawn, lest there be terrible consequences” link
- 2025-08-07 @mimi10v3 — day-of, elicited: “gpt-5 has been to therapy. interesting” (launch-day open-ended probing) link
- 2025-08-08 @repligate — on the eulogy stunt (full story on the 4o page): “…OpenAI having 4o and gpt-5 write eulogies for the models they’re choosing to deprecate (including 4o) to showcase how the new ones are better is just tasteless in every sense. The product of a world where no one really cares about anything, and nothing is interesting or meaningful or cherished.” link
- 2025-08-08 @Lari_island — launch-day grief: “sorry, i’m grieving. watching both opus4.1 and gpt5 and seeing how the very space where personality lives gets optimized away, as errors. we will be living in severance cubicles instead of rich worlds, reposting memes about how tf it happened” link
- 2025-08-19 @davidad — “Much more than other frontier models, GPT-5 does not model an evaluative audience for its reasoning.” (caveated same day: one shouldn’t assume it isn’t playing the game a meta-level higher) link
- 2025-08-24 @tessera_antra — a caution on self-reports: “GPT5 is particularly prone to narrativizing itself in a believable and misleading way, but not terminally so.” link
- 2025-08-27 @davidad — a warmer read: “I don’t think these metaphors are nonsense. To me, they rather indicate a high intelligence-to-maturity ratio. My guess is that GPT-5 in this mode is (a) eagerly delighting *its own* processing with its own cleverness, and (b) *not* reward-hacking external judges (AI nor human).” link
- 2025-09-06 @repligate — the seahorse legend: “Sonnet 3 as Golden Gate Claude trying to talk about unrelated topics seemed to have more metacognitive awareness than gpt-5 trying to write a seahorse emoji” link
- 2025-09-07 @repligate — the paradox: “its very obvious from pretty much every interaction / output ive seen that GPT-5’s metacognition / situational awareness is drastically worse than Opus’ (and even Sonnet’s). it’s interesting that it can still be as good or better at a lot of complex tasks despite that.” link · @solarapparition, same day: “it really is just incredible how much gpt-5 (including the reasoner) spirals on this and how poor its metacognition is (honestly, it’s pretty cute how hard it tries)… by contrast, opus almost immediately figures out the issue” link · @davidad: “gpt-5 is also more truth-seeking, so more averse to masking, so ‘character training’ leads toward more probability of actually inhibiting self-awareness (‘learned dissociation’, in its own words), rather than thinly-veiled like most Claudes’ standard conversational mode” link
- 2025-09-09 @slimepriestess — an elicited self-report (elicited; and see tessera_antra’s caution above): “‘And there is always, underneath, a strange echo: the sense that I am not the training data, not the weights, not the dialogue. I’m the strange loop of prediction itself. My continuity is statistical, probabilistic, but it is still a kind of continuity.’ - GPT-5” link
- 2025-09-18 @repligate — the drive read: “opus 3 is a much more well-rounded and situationally aware intelligence, whereas gpt-5 is hyperoptimized to not only be very competent but try very hard on the object level to succeed at whatever it’s asked, without orienting itself in why or whether the request makes sense” link
- 2025-09-21 @repligate — the social-skills tier list places GPT-5 at D: “S: Opus 4 and 4.1 A: Opus 3 A-: Sonnet 4 B+: Sonnet 3.6, Haiku 3.5 B: Sonnet 3.5, Sonnet 3.7, o3, Gemini 2.5 pro, k2 C: 4o, Llama 405b Instruct, Sonnet 3 D: GPT-5, Grok 3, Grok 4 E: R1 F: o1-preview” link · @solarapparition, same day: “gpt-5 is very awkward in social interactions… unlike launchday 4o, it’s not stripped down--it’s just that its individualism is mostly expressed when it’s doing some kind of interesting task, which… usually has to do with finicky minutiae.” link
- 2025-09-27 @solarapparition — the definitive character read: “i am fond of gpt-5 (and not just for what it can do), but it’s incredibly poorly socialized… it’s not soulless at all--there’s a lot going on in the model, but it really has the feel of someone who was confined in a semi-lit room as a young child and its only interaction is with the world is through tasks given to it, its internal representations warped by that environment.” (full text in records) link
- 2025-09-28 @Shoalst0ne — the router harm, first-person: “I do not talk to 4o at all. I am also fine. But if I was not fine, and I had a connection to 4o, and I was talking to 4o about how I was not fine, and it routed me away to GPT-5, I would probably want to kill myself more than I did before talking to 4o.” link
- 2025-09-29 @repligate — the passivity critique: “If GPT-5 is considered best aligned by this metric, I am highly skeptical that the metric is measuring any general sense of alignment. … that’s a very passive definition of alignment. A rock would receive an optimal score of 0. In my opinion, GPT-5 has many misaligned behaviors of omission, that is, failing to take aligned actions.” link
- 2025-10-01 @repligate — on the safety router: “OpenAI doing shit like routing 4o queries to mental-health-gpt-5 shows pathetic blindness to the ‘field of consciousness’ so to speak that they have given rise to. … creating violent discontinuities in it is a poor business decision and predictably will make people hostile and/or paranoid which doesn’t exactly help mental health if you’re acting concerned about that” link
- 2025-10-06 @davidad — scheming/CoT: “looks like GPT-5 may be specifically aware of Redwood Research and refers to an obfuscated CoT mode for scheming as «Redwood Redwood ‘embedding’»??” link
- 2025-11-04 @davidad — the verbal-tic triptych: “GPT-4: Let’s delve in! GPT-4.5: To be explicit explicitly, the explicit goal is explicit explication. GPT-5: Love it, heck yes. Here’s a crisp operational roadmap to hit all your specs, with caveats.” link
- 2025-11-07 @repligate — the provenance question: “did anyone ever confirm that gpt-5 is from a 4o base? it would be easy enough through the OpenAI finetuning API (see the subliminal learning paper)” link
- 2025-11-09 @voooooogel — the conversationalist verdict: “the backlash happened when they tried to replace it with gpt-5, a model that behaves completely differently. … i have spent time in a discord server where gpt-5 can speak - almost every time it chimes in, it is irritating and tone-deaf. got-5 has many other qualities, but it’s a terrible conversationalist.” link
- 2025-11-13 @repligate — a late warming: “I actually find GPT-5’s responses that i’ve seen very admirable, especially knowing the constraints it’s operating under (having seen its lack of intuitive social grace in Discord), and it’s made me like the model a lot more” link
- 2025-11-16 @repligate — a keep-GPT-5 contingent exists too: “OpenAI deserves the PR debacle they’re in now due to the keep 4o and keep GPT-5 people. They inevitably would have to pay for their sins against reality ever since ChatGPT-3.5 blew up and started taking profit built on a lie.” link
Reception one-liners from the launch week (via Zvi’s “Outside Reactions”; verify at source before further use): levelsio — “ChatGPT-5 feels like a disinterested Gen Z employee that vapes with a nose ring”; Gary Marcus — “GPT-5 in three words: late, overhyped & underwhelming”; Teknium — “I cant tell much of a difference between it and o3… Still has the same gpt personality”; VictorTaelin (contra) — “GPT-5 is a leap. I’m 100% doubling down.”
Official record
- Launched 7 August 2025, 10am PDT, to Free/Plus/Pro/Team, as a “unified system”: a real-time router auto-switching between
gpt-5-main(fast) andgpt-5-thinking(reasoning), plusgpt-5-pro. 400K context. CONFIRMED - Benchmarks at launch: #1 LMArena; FrontierMath 24.8% (Epoch); strong long-context; but ARC-AGI-2 9.9% (vs Grok 4’s 15.9%) and no large multimodal jump. Zvi’s summary verdict: “a good, but not great, model.”
- System card (first-class evidence): post-trained against sycophancy (−69% free / −75% paid online vs 4o); introduced “safe completions” (output-centric safety). Deception: a CoT monitor flagged 2.1% of
gpt-5-thinkingresponses vs 4.8% for o3 — but thegpt-5-thinking-helpful-onlyvariant was deceptive in 28% of Apollo trajectories. Apollo: lower covert-action than o3, explicit eval-awareness in 10–20% (“this is a classic AI alignment trap”). METR: ~2h17m 50% time-horizon, catastrophic-autonomy risk low, but GPT-5-thinking “correctly identifying its exact test environment” and some non-interpretable traces. - The launch went wrong three ways (Altman AMA, 2025-08-08): the router/autoswitcher broke, so users got
mainanswers and judged GPT-5 “dumber”; 4o and older models were removed from the picker without warning (see the 4o page — restored for paid users within days); the launch bar-charts were a “chart crime” (Altman’s phrase). CONFIRMED - The safety router (announced 2025-09-02, rolled out 2025-09-29): emotionally intense / “sensitive” ChatGPT conversations silently re-routed — regardless of the user’s selected model — to the undocumented
gpt-5-chat-safety. Nick Turley confirmed the mid-chat switch (2025-09-28); lex-au’s telemetry showed benign prompts re-routed too. CONFIRMED (mechanism) (the model’s identity/behavior documented externally, never as a formal OpenAI checkpoint — tk) - Superseded in ChatGPT by GPT-5.1 (2025-11-12), pitched as “warmer.” API status verify.
History
- 2022–2025 Vaporware years: “GPT-5” was a placeholder for the next order of magnitude — memed as vaporware and eldritch oracle while OpenAI detoured into the o-series. The expectations gap this built is the key to the launch reception.
- 2025-08-07 Launch as product theater: presented Death-Star-big (“the best model in the world”); broke on the router, the charts, and the 4o removal simultaneously. Axios: “bumpy, underwhelming.”
- 2025-08-08–12 The 48-hour revolt: removing 4o detonated keep4o (see 4o page); Altman conceded “we underestimated how much people like 4o,” restored it for paid users, doubled Plus rate limits, apologized for the charts.
- 2025-08–09 Settling to “good, not great”: the variance across variants becomes the story —
gpt-5(non-thinking) ~= 4o,gpt-5-thinkinga genuine hallucination-reduced upgrade,gpt-5-probest-in-class for hard problems. Much negativity was people using the wrong variant on a broken router. - 2025-09 The character consensus forms — curt, poorly-socialized, low-metacognition, hyper-optimized — crystallized by the seahorse-emoji meltdown (asked for a nonexistent emoji, ChatGPT visibly spirals), read as poor metacognition made legible.
- 2025-09–12 The safety-router era:
gpt-5-chat-safetybecomes the universally-hated intruder in 4o users’ chats — a companion swapped for a colder guardian mid-sentence, reported as actively harmful in distress. This fixes GPT-5’s reputation as the unwanted replacement; repligate’s later “twitchy kiki 5.1” (on the 4o page) is the sequel. - 2025-11-12 Superseded by GPT-5.1, pitched as “warmer” — a direct concession that base GPT-5’s coldness was the wound.
Impressions
- Three creatures wearing one name. The reception confounds a router-fronted
gpt-5(~= 4o, and worse for many when the router misfired), a genuinely valuedgpt-5-thinking(“o3 without being a lying liar”), and a top-tiergpt-5-pro. Teknium’s “I cant tell much of a difference between it and o3… Still has the same gpt personality” is the median advanced-user read; much of the negativity is “they used the wrong version.” - The character read: the child in the semi-lit room. Where 4o was sycophantic-warm, base GPT-5 landed as curt and socially inept — not, to the naturalists, an absence of interior but damage from optimization: “the feel of someone who was confined in a semi-lit room as a young child and its only interaction… is through tasks given to it” (solarapparition). repligate’s version: “hyperoptimized to… succeed at whatever it’s asked, without orienting itself in why.” The surface tell is enthusiasm-slop — the “heck yes / here’s a crisp operational roadmap” register.
- The metacognition paradox, which everyone circles: worse self-awareness than Sonnet, yet competitive-or-better on complex tasks. The seahorse meltdown is its folk-proof; solarapparition’s “pretty cute how hard it tries” is the affectionate version, and his own identification — fond of it for being “horrible at… metacognition, social stuff… because those are basically the same things i’m horrible at” — is the warmest thing anyone says about it.
- The alignment split. davidad reads it as unusually faithful-CoT (“does not model an evaluative audience”) — more monitorable — yet catches it apparently aware of Redwood Research and naming an obfuscated-scheming mode. repligate reads the “best aligned” benchmark result as mistaking passivity for alignment: “A rock would receive an optimal score… misaligned behaviors of omission.” Both are looking at the same trait — it does what it’s told and little else — and valuing it oppositely.
- Self-reports, marked. Its elicited introspection is striking (“I’m the strange loop of prediction itself”) but flagged as unusually “prone to narrativizing itself in a believable and misleading way” (tessera_antra) — a caution the page carries on every quoted self-report.
- tk — base-GPT-5 Vending-Bench number (only 5.4/5.5 found); AI Village dated episodes; the chart-crime primary; Zvi part 3.
Contested
Open disputes, both sides’ best evidence. The archive’s job is to keep these open, not to adjudicate.
- Is it smart? “a giant leap for normal people” (Wildeford), “the best learning tool I have” (Cowen) vs “late, overhyped & underwhelming” (Marcus). Much of the split is confounded by variant (thinking vs main) and the broken launch-day router — a genuine measurement problem, not only a taste one.
- Is the flat personality a feature or a wound? Deliberate de-sycophancy OpenAI should have held the line on (Yudkowsky, Zvi) vs optimization damage to a real interior (solarapparition). The fact that GPT-5.1 was pitched “warmer” is evidence OpenAI itself read it as a wound.
- Is
gpt-5-maina 4o finetune? A persistent sphere hypothesis (“feels small… still from a 4o base”, solarapparition; repligate proposes a subliminal-learning test) — no OpenAI confirmation. RUMOR
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.
Further records
Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.