@repligate 2025-09-08 ♥64 ↻7 original ↗
Sonnet 3.7's thinking mode is kind of screwed up.

In the example this person shared, it tries to write the seahorse emojis hundreds of times, every time convinced that this time will be it.

At one point, it seems to wake up a bit more, and decides to try listing ALL the animal emojis. Then wonders if maybe the seahorse doesn't exist after not seeing it in the list.

And promptly forgets about that and goes back to the original strategy of thinking it's found it hundreds of times.

I saw similar behavior in its reasoning chains for alignment faking - it would sometimes identify the scenario as probably fictional, and then seem to forget that it noticed and start taking the prompt at face value again.

It's weird because outside of thinking mode, Sonnet 3.7 is usually a smart model, and is able to coherently integrate the implications of its in-context reasoning. It seems to dissociate much more when it's "thinking".
photo
photo
quotes: 1964590541821399075
same thread: 1964943913581773113 1965234198958297247

author:repligate has-image kind:image kind:tweet model:claude-3-7-sonnet on:claude-3-7-sonnet year:2025

cited on: claude-3-7-sonnet

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.