@QiaochuYuan 2026-06-16 ♥104 ↻0 original ↗
there's a bunch of questions i was asking (eg media analysis questions, "speculate on the meaning of this movie") where gpt-5.5 and opus 4.8 say things that sound reasonable but if you pay closer attention or double-check against other sources you'll see where they cut a lot of corners and are somewhat bullshitting based on superficial details. fable did a lot less of this and a lot more stuff that seemed like it actually held up and was getting to the heart of things. it seemed meaningfully better at relevance realization. of course i didn't have enough time to really thoroughly test this

author:qiaochuyuan kind:tweet model:claude-opus-4-8 model:fable model:gpt-5-5 on:claude-opus-4-8 on:gpt-5-5 year:2026

cited on: claude-opus-4-8 · gpt-5-5

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.