@voooooogel 2024-12-21 ♥21 ↻0 original ↗
@fchollet "high efficiency" (less compute) is 33M tokens at 6 samples. "low efficiency" (more compute) is 5.7B tokens at 1024 samples. 5.7B / 1024 * 6 = ... 33M. so is the 176x compute scaling just increased majority voting sample count, not longer reasoning? huh?
https://t.co/qw6Wh4vYAa
in reply to: 1870318275676233858
same thread: 1870315002172952650 1870318275676233858

author:voooooogel kind:tweet thread-context year:2024

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.