@kromem2dot0 yeah, which is why i think a METR-like "lowest common denominator environment" benchmark is ~fine, as long as you don't overindex on the y values or model-to-model comparisons. the shape of the curve seems to be accurate
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.