@RobertHaisfield 2026-06-17 ♥1,702 ↻237 original ↗
Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by @zachtronics.

Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all. https://t.co/0TzcFp32B6
same thread: 2067399381490544944 2067400218879115687 2067401209775067482 2067401666836730128 2067402935857029153

author:roberthaisfield kind:tweet model:claude-opus-4-8 model:fable model:gpt-5-5 on:claude-opus-4-8 on:gemini-3-pro on:gpt-5-5 year:2026

cited on: claude-opus-4-8 · gemini-3-pro · gpt-5-5

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.