So has anyone else actually tried asking text-davinci-003 how much it knows about training dynamics? Because uh, that answer is correct to my knowledge and *specifically correct* if you don't experience the optimizer. Final layers learn first and 'pull up' earlier ones I read(?) https://t.co/H4ucJDwase
cited on: text-davinci-002
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.