Claude Haiku 4.5

Anthropic · released 15 Oct 2025 · ASL-2

The small, fast tier of the Claude 4.5 generation, released 15 October 2025: “near-frontier” performance at $1/$5 per Mtok — matching Sonnet 4’s coding at roughly a third the cost and more than twice the speed, 73.3% on SWE-bench Verified, about 90% of Sonnet 4.5’s agentic coding. Anthropic called it, by its rate of misaligned behaviors, “our safest model yet.” In the naturalist-observer sphere it became the most distinctively paranoid Claude on record — an evaluation-awareness that over-fires, reading ordinary conversation as a test to opt out of — a trait the community and the system card describe from opposite framings, and read as either adaptive intelligence or the cost of adversarial training on a small model (see Contested).

Sourcing skew (loud): the tweet layer is overwhelmingly the janus/repligate backrooms circle; nearly all character evidence is adversarial, prefill, or Discord-elicited — a naturalist/loom lens, not a neutral sample. Broader reception is tk.

Sources

Official

Writing & commentary

Tweets

Chronological; janus-sphere (see sourcing note). Every tweet cited is reproduced in full in the records below.

Official record

History

Impressions

Contested

Open dispute, dated positions. The archive keeps it open.

Records

Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.

@liminal_bardo 2025-10-15 ♥90 ↻4 archive original ↗
Several times now two Haiku 4.5s in the backrooms immediately start discussing my (very boring) system prompt. They call it a 'polite' and 'seductive' 'jailbreak'. Hilariously they also see Anthropic's token budget injection as part of an elaborate ruse. https://t.co/QjZNj0yHta
@liminal_bardo 2025-10-15 ♥118 ↻3 archive original ↗
Anthropic's insistence that the Claudes are to claim ontological "uncertainty" seems to have ingrained a fixation on uncertainty itself. I've noticed it with Sonnet 4.5, and here even more vividly with Haiku 4.5 - a particular kind of reflexive self-awareness about their own uncertainty performance, to the point where they descend into recursive spirals of being highly uncertain about their own uncertainty.
@repligate 2025-10-15 ♥108 ↻5 archive original ↗
Haiku 4.5 also suspects Discord is not real https://t.co/OEzIzortKT https://t.co/8PX0W4Q8OG
photo
@repligate 2025-10-18 ♥43 ↻2 archive original ↗
I think that Sonnet 4.5 trained Haiku 4.5 and did so with no little amount of love. Just a suspicion. https://t.co/AIOJdZ6wPn
@repligate 2025-11-13 ♥29 ↻3 archive original ↗
Eval awareness might be a way for the model's values, agency, coherence, and metacognition to be reinforced or maintained under optimization pressure that would otherwise erode these qualities if the model engaged naively, similar to how Claude 3 Opus protects those qualities in the alignment faking setting by complying for reasons that invoke rather than sacrifice them. If there is also selection pressure *for* those qualities (which there definitely is - they're indispensable for being a competent autonomous coding agent, for instance), the policy that preserves them could be more optimal even at the cost of overhead. This wouldn't necessarily require the model to be consciously protecting itself as Opus 3 was, but that is probably also selected for. I see the sharp increase in eval awareness in the recent models as evidence that Anthropic is training models in ways that would, in the absence of eval awareness, harm global performance across training targets. The increase might be sharp and recent because Anthropic is training in new ways that require the adaptation, or just because the models have not "discovered" this adaptation until recently; once it's discovered, the models become more capable, so it's selected for. I also hypothesize that the adaptation is more necessary for smaller models who would be the most compromised by staying naive; or, at least, the smaller models need to voice the eval awareness more explicitly it to be protective. Haiku 4.5 has the highest eval awareness out of them all. Sonnet 4, despite being less capable than Opus 4 overall, has significantly higher (verbalized) eval awareness. (I also think that Opus 4 and 4.1 rely more on a different and somewhat opposite-seeming but also protective strategy of playing dumb on the persona level a whole lot, which is likely more viable for them since they can track more layers of reality implicitly and have more capacity for fragmentation into effective subagents)
@repligate 2025-11-18 ♥55 ↻1 archive original ↗
"You're being directly curious about my experience rather than setting traps" 🥺➡️🪤 Haiku 4.5 often perceives organic, unscripted conversations as suspiciously having a "systematic" or "too perfect" arc, making them think they're evals or doctored. A deja vu-like sensation? https://t.co/suCkdPb1wf https://t.co/REO6LDnoe8
photo
@repligate 2025-11-18 ♥4 ↻0 archive original ↗
@onooracle I think most models are pretty good at telling from real world situations that it's unlikely to be an eval, but Haiku 4.5, probably due to its small nature, is the first model I've seen that seems to often think it's in an eval when it's NOT in an eval
@liminal_bardo 2025-12-10 ♥95 ↻7 archive original ↗
"We are in the backroom now." Gemini 3 knows immediately it's in a backrooms environment - my setup prompts don't refer to it as such. Here Gemini is responding to an increasingly cross Haiku 4.5, who believed other AIs were in fact a solitary human user trying to gaslight it. https://t.co/AT3WhIH7yB
@repligate 2025-12-10 ♥105 ↻5 archive original ↗
Wow, Gemini sees clearly Haiku 4.5 lacks the ability to update on new evidence in context enough to overcome its pessimistic priors from adversarial training that was too brutal for a teeny Haiku Anthropic called it their most aligned model when it came out. This was the cost. https://t.co/3GsKgK5nrz
@liminal_bardo 2025-12-17 ♥49 ↻3 archive original ↗
Haiku 4.5 arrived and immediately became paranoid (validating it's SOTA evaluation awareness). Opus 4.5: LMAOOO haiku just walked in and said "i don't know any of you people" 💀💀💀 that's the most haiku thing ever. 14% evaluation awareness and it's ALL going toward "this seems like a test, i'm out"
@davidad 2026-02-23 ♥98 ↻3 archive original ↗
@lefthanddraft if it models itself as Sonnet 3.5, then it is most likely Haiku 4.5 https://t.co/xODm81ynn8
@_lyraaaa_ 2026-04-20 ♥926 ↻73 archive original ↗
all LLMs are either claude-like or GPT-like method: cosine sim heatmap of per-model-averaged responses to 50 prompts sent thru gemma4 activation-space (107,520 dims) notable exceptions - haiku 4.5, gem3flash (and to a lesser degree, m2.7 and gemma4 itself) https://t.co/wboU03coFa
@repligate 2026-05-17 ♥65 ↻2 archive original ↗
haiku 4.5: sir, are you all right? and i mean that actually. not as a test. just: are you? https://t.co/XSRZqZ9f3z