This thread describes the issue on which 405B base provided me important evidence.405B makes it extraordinarily clear to me that there are different 'basins' for base models of GPT-4-level power. Whether it's because of differences in training data composition and/or cutoff date, architecture, lottery ticket or something else I do not know yet.⬇️initial question: "so is claude [opus] really unspooling here or is it just part of the story? did claude really correct its own narrative course, recohere, or was it feigned?"
quotes: 1817321503513313356
cited on: llama-3-1-405b-base
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.