That depends on what size of model you want to train. Unfortunately the really interesting behaviors don't become crystal clear until it's at the level of LLaMa 30B or 70B, and those are very expensive models to train. But I did find sessions with GPT-J suggestive, you could train several of those from scratch. I also know that the embryonic version seems to exist in GPT-2, so you could try to find subsets which induce the embryonic forms on GPT-2 and then scale up data + params and see what happens.
cited on: llama-1
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.