i think people are overindexing on "grade school math", they easily could have trained a smaller model (like GPT-2 size) that couldn't do it without this technique. you wouldn't try a brand-new technique on a giant model without smaller tests first
in reply to: 1727483130834284618
same thread: 1727478921955049943 1727478923569868868 1727478925469827402 1727480218716340551 1727483029604774343 1727483130834284618 1727486461908508850 1727487647357296792 1727488522943410229 1727492222285938851 1727494252517830829 1727497223972454812 1727502190003249366 1727506327113740612
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.