@voooooogel 2025-02-01 ♥8 ↻0 original ↗
@max_paperclips i think the ideal would be to seed a few structures and then hope R1-Zero style that the model can generalize beyond that, into some kind of "meta-reasoning". the humanizing SFT they did to R1 seems to have actually clamped down on that creativity, sadly.

author:voooooogel kind:tweet model:deepseek-r1 model:deepseek-r1-zero on:deepseek-r1-zero year:2025

cited on: deepseek-r1-zero

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.