@xlr8harder 2025-02-28 ♥86 ↻1 original ↗
This is my new conspiracy theory btw. The reason we don't have a benchmark-maxxed GPT-4.5 is the same reason we don't have an Opus 3.5: benchmark-maxxing extinguishes the difficult-to-measure big model smell, and so the extra large model provides limited benefit after tuning.

author:xlr8harder kind:tweet model:claude-3-opus model:gpt-4-5 on:claude-3-5-opus on:gpt-4-5 year:2025

cited on: claude-3-5-opus · gpt-4-5

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.