@KeyTryer I'm not sure what "as expected" means - in terms of pretraining loss, probably - but the expectation should be that emergent phenomena that you can't predict in advance emerge with unprecedented scale. This was very true for GPT-4 sized vs GPT-3 sized models
same thread: 1963439548383592569 1963440322714771631 1963441257780228180 1963442555896697190 1963443052133191996 1963444433804038339 1963445092745973858
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.