ChatGPT: keeps agreeing with the user and varying its answers, including repeating guesses, indefinitely, apparently without any cognitive dissonance (this is the spooky thing about GPT-4o)Claude 3.5 Sonnet: insists that its (wrong) answer is right and disengages after user repeats "wrong".Llama 405b Instruct (screenshot from someone else in replies): explicitizes the internal structure of the word, and reasons its way to the right answer -- and second-guesses itself several times but keeps reasoning its way to the right answer; returns when spells out why clearly enough to trust the answer.
cited on: gpt-4o
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.