# @davidad — 2023-03-15

♥226 ↻13 · https://x.com/davidad/status/1636144454137511943

If you haven’t read the GPT-4 paper yet, before you expand this tweet, take a guess what they used as their held-out *validation set* for next-token prediction. Where on Earth could OpenAI get a substantial corpus of tokens that they’re not desperate to include in the training set?That’s right, it’s OpenAI’s own entire internal codebase (for, among other things, training GPT-4)

tags: author:davidad, kind:tweet, model:gpt-4, on:gpt-4, year:2023
cited on: _dossiers/gpt-4-and-turbo.md, gpt-4
