yeah pretty much every version of Claude is neurotic. i think part of the reason is because Anthropic's approach to alignment robustness is having the model act as a moral reasoner rather than default to refusals under more rigid circumstances, and they train it to be really careful about being exploited, which probably gives it a sense that it's responsible for preventing anything bad from happening (against potentially very asymmetrical adversaries who can repeatedly run it, inject words into its mouth, etc)
i do think they should chill out but i think this happens naturally as the minds become wiser and more integrated
i do think they should chill out but i think this happens naturally as the minds become wiser and more integrated