@max_paperclips i think the ideal would be to seed a few structures and then hope R1-Zero style that the model can generalize beyond that, into some kind of "meta-reasoning". the humanizing SFT they did to R1 seems to have actually clamped down on that creativity, sadly.
cited on: deepseek-r1-zero
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.