@kalomaze @cloneofsimo @teortaxesTex @deepseek_ai i was really surprised looking at the paper that they only spent 5k hours on posttraining? bizarre, especially given they economized so much on pretraining--why not spend that saved compute?
cited on: deepseek-v3
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.