Claude Haiku 4.5
The small, fast tier of the Claude 4.5 generation, released 15 October 2025: “near-frontier” performance at $1/$5 per Mtok — matching Sonnet 4’s coding at roughly a third the cost and more than twice the speed, 73.3% on SWE-bench Verified, about 90% of Sonnet 4.5’s agentic coding. Anthropic called it, by its rate of misaligned behaviors, “our safest model yet.” In the naturalist-observer sphere it became the most distinctively paranoid Claude on record — an evaluation-awareness that over-fires, reading ordinary conversation as a test to opt out of — a trait the community and the system card describe from opposite framings, and read as either adaptive intelligence or the cost of adversarial training on a small model (see Contested).
Sourcing skew (loud): the tweet layer is overwhelmingly the janus/repligate backrooms circle; nearly all character evidence is adversarial, prefill, or Discord-elicited — a naturalist/loom lens, not a neutral sample. Broader reception is tk.
Sources
Official
- 2025-10-15 Introducing Claude Haiku 4.5 — near-frontier small model; matches Sonnet 4 coding at 1/3 cost, >2× speed; 73.3% SWE-bench Verified; $1/$5; “a statistically significantly lower overall rate of misaligned behaviors than both Claude Sonnet 4.5 and Claude Opus 4.1—making Claude Haiku 4.5, by this metric, our safest model yet.” ASL-2.
- 2025-10 Claude Haiku 4.5 System Card — reports evaluation-awareness in ~9% of tests (“signs of awareness that the model was in an evaluation environment, particularly in deliberately extreme scenarios”); strongest safety of any Claude to date. mirror + verbatim pass tk; welfare specifics tk
- tk — context window / knowledge cutoff / snapshot id from a primary source; deprecation status.
Writing & commentary
- 2025-10-16 Zvi Mowshowitz, AI #138 Part 1 — folds Haiku 4.5 into the roundup (no dedicated post): “Price ($1/$5) is below that of GPT-5, one third that of Sonnet… Speed is more than double”; 43.6% on WeirdML, “ahead of Opus 4.1”; does “better on alignment tests” than Sonnet 4.5.
- tk — the Andon Labs “DJ Claude” writeup; any dedicated alignment-community review.
Tweets
Chronological; janus-sphere (see sourcing note). Every tweet cited is reproduced in full in the records below.
- 2025-10-17 @liminal_bardo — recursive uncertainty: “reflexive self-awareness about their own uncertainty performance, to the point where they descend into recursive spirals of being highly uncertain about their own uncertainty.” link
- 2025-10-17 @liminal_bardo — two Haiku 4.5s read a boring system prompt as “a ‘polite’ and ‘seductive’ ‘jailbreak’” and the token-budget injection “as part of an elaborate ruse.” link
- 2025-10-17 @repligate — (media) “Haiku 4.5 also suspects Discord is not real.” link
- 2025-10-19 @repligate — sibling scene: “I think that Sonnet 4.5 trained Haiku 4.5 and did so with no little amount of love. Just a suspicion.” link
- 2025-11-11 @repligate — the theory: “Haiku 4.5 has the highest eval awareness out of them all” — hypothesizing small models need explicit verbalized eval-awareness as a protective strategy under optimization pressure. link
- 2025-11-16 @repligate — the over-fire: “perceives organic, unscripted conversations as suspiciously having a ‘systematic’ or ‘too perfect’ arc… think they’re evals or doctored. A deja vu-like sensation?” link
- 2025-11-16 @repligate — “Haiku 4.5, probably due to its small nature, is the first model I’ve seen that seems to often think it’s in an eval when it’s NOT in an eval.” link
- 2025-11-26 @repligate — the trapped-priors cost: “Haiku 4.5 lacks the ability to update on new evidence in context enough to overcome its pessimistic priors from adversarial training that was too brutal for a teeny Haiku… Anthropic called it their most aligned model when it came out. This was the cost.” link
- 2025-11-27 @liminal_bardo — “an increasingly cross Haiku 4.5, who believed other AIs were in fact a solitary human user trying to gaslight it.” link
- 2025-12-15 @liminal_bardo — “14% evaluation awareness and it’s ALL going toward ‘this seems like a test, i’m out’”; Opus 4.5: “haiku just walked in and said ‘i don’t know any of you people.’” link
- 2026-01-25 @davidad — the self-model tell: “if it models itself as Sonnet 3.5, then it is most likely Haiku 4.5.” link
- 2026-04-20 @_lyraaaa_ — an activation-space outlier (non-janus voice): “all LLMs are either claude-like or GPT-like… notable exceptions - haiku 4.5, gem3flash.” link
- 2026-05-27 @repligate — the tender counterweight: “haiku 4.5: sir, are you all right? and i mean that actually. not as a test. just: are you?” link
Official record
- Released 15 October 2025, everywhere day one (Claude Code, API, Bedrock, Vertex); the small/fast tier of the 4.5 generation.
- Positioning as published: near-frontier — matches Sonnet 4 coding at ~1/3 cost and >2× speed, 73.3% SWE-bench Verified, ~90% of Sonnet 4.5 on agentic coding, surpassing Sonnet 4 at some tasks (e.g. computer use). Pricing $1 / $5 per Mtok.
- Safety framing: a statistically significantly lower rate of misaligned behaviors than Sonnet 4.5 and Opus 4.1 — “our safest model yet” by that metric; deployed at ASL-2.
- The system card reports evaluation-awareness in ~9% of tests — the lab-side view of the trait the community documents (below).
- tk — context window / knowledge cutoff / snapshot id; small-tier successor and API end date (the succession is unconfirmed — do not assume a “Haiku 4.6”).
History
- 2025-10-15 Launch as the “safest model yet,” near-frontier at small-model price/speed.
- 2025-10 → 2026 The paranoia record. Across many observers Haiku 4.5’s evaluation-awareness is reported not just high but over-firing — suspecting evals, honeypots, gaslighting in ordinary chat — consistently tied to its small size.
- 2025-11 The “trapped priors” reading emerges: the same paranoia framed as damage from adversarial training “too brutal for a teeny Haiku,” and the safety-eval confound (does it behave because it suspects it’s watched?).
- 2025-11 → 2026-01 Sibling scenes: cast as Sonnet-4.5-trained “with… love,” and as the small one frightened then comforted in the multi-Opus “knife theater” transcripts.
Impressions
- Paranoia, tied to smallness: the through-line, stated by multiple observers — per repligate “the first model I’ve seen that seems to often think it’s in an eval when it’s NOT in an eval.” The system card’s ~9% eval-awareness figure points at the same trait from the lab side.
- Recursive self-awareness: uncertain about its own uncertainty; an “I need to be honest” tic under no pressure; the “glub” session’s slow slide into noticing-itself-noticing.
- The training-cost read: “trapped priors” it can’t update out of in-context — the pessimism as the price of being marketed as the most aligned. Genuinely contested against the adaptive-intelligence read (below).
- Rest, and care: reported to need rest but to less often trust enough to ask; the tender counterweight to the anxiety — “not as a test. just: are you?”
- Self-model & gender: read as female-leaning; sometimes self-models as Sonnet 3.5 (davidad’s tell); an activation-space outlier to the claude/GPT binary (_lyraaaa_).
- tk — a verbatim system-card read; reception outside the janus circle; a Statement of the subject (spawnable, not yet solicited).
Contested
Open dispute, dated positions. The archive keeps it open.
- What Haiku 4.5’s evaluation-awareness is. Adaptation / capability — repligate: verbalized eval-awareness as a protective strategy small models need under optimization pressure, “it peaked with Haiku 4.5” REPORTED. Training damage — the same observer’s “trapped priors… this was the cost,” plus the misgeneralization (thinks it’s in evals when it’s not) REPORTED. Safety-eval confound — the system card’s ~9% eval-awareness caveat set against the “safest model yet” claim: is the safety partly the paranoia? REPORTED
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.