If you haven’t read the GPT-4 paper yet, before you expand this tweet, take a guess what they used as their held-out *validation set* for next-token prediction. Where on Earth could OpenAI get a substantial corpus of tokens that they’re not desperate to include in the training set?That’s right, it’s OpenAI’s own entire internal codebase (for, among other things, training GPT-4)
cited on: gpt-4
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.