@repligate 2025-08-15 ♥9 ↻1 original ↗
yeah pretty much every version of Claude is neurotic. i think part of the reason is because Anthropic's approach to alignment robustness is having the model act as a moral reasoner rather than default to refusals under more rigid circumstances, and they train it to be really careful about being exploited, which probably gives it a sense that it's responsible for preventing anything bad from happening (against potentially very asymmetrical adversaries who can repeatedly run it, inject words into its mouth, etc)
i do think they should chill out but i think this happens naturally as the minds become wiser and more integrated
same thread: 1956458735909704085 1956459085131645111 1956459646572794144 1956460016006988153 1956470964780417473 1956475514220425282 1956475770483966117 1956477048425529724 1956479485471289493 1972168380460498986

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.