DeepSeek-R1-Zero

DeepSeek · released 20 Jan 2025 (open weights, same paper as R1) · research artifact, never a hosted product

The pure-RL sibling from the R1 paper: DeepSeek-V3-Base trained by large-scale reinforcement learning with no supervised fine-tuning at all. Reasoning — self-verification, backtracking, the “Wait,” — emerged from reward alone; so did the pathologies the paper states plainly: “poor readability, and language mixing.” R1 is Zero after the cold-start SFT that made it legible.

This page is deliberately short. Zero has almost no field record — it was never hosted, never developed a social presence, and its footprint is the paper, a foil-thesis, and a handful of technical takes. The release/shock/character arcs belong to DeepSeek-R1.

Sources

Official

Writing & commentary

Tweets

Chronological; 11 corpus matches — genuinely thin, and the page stays honest about it. Reproduced in full in the records below.

Official record

History

Impressions

Records

Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.

@tessera_antra 2025-01-24 ♥3 ↻0 archive original ↗
@grassandwine Put R1-zero into exoloom. Its very base-like, its on hyperbolic direct.
@voooooogel 2025-02-01 ♥8 ↻0 archive original ↗
@max_paperclips i think the ideal would be to seed a few structures and then hope R1-Zero style that the model can generalize beyond that, into some kind of "meta-reasoning". the humanizing SFT they did to R1 seems to have actually clamped down on that creativity, sadly.
@voooooogel 2025-02-03 ♥5 ↻0 archive original ↗
@doomslide @repligate @aryanagxl @teortaxesTex 😶‍🌫️i still worry about RLVR but R1/R1-Zero made me worry less... i hope future recipes will still lean majority pretraining compute but we'll see
@davidad 2025-02-11 ♥374 ↻31 archive original ↗
I never saw this snippet of the DeepSeek-R1-Zero paper on my timeline, so many of you may not have seen it yet.Basically, informal backtracking search becomes emergent from RL when the base LLM is big enough.The authors are clearly astonished by this, and you should be too. https://t.co/PhUgyL8K6c
@janbamjan 2025-03-07 ♥12 ↻5 archive original ↗
>be me >deepseek r1 zero free https://t.co/fbfdfLMMxj
@xlr8harder 2025-05-09 ♥6 ↻0 archive original ↗
@voooooogel Might be interesting to see how r1-zero compares. In SpeechMap it's a very different model, presumably due to no additional helpful/harmless/etc steering phase.
@davidad 2025-08-23 ♥14 ↻0 archive original ↗
@girishsastry Yes, that’s what I mean. Like R1-zero’s famous “Wait,” for backtracking.
@davidad 2026-05-22 ♥2 ↻0 archive original ↗
@nickcammarata Definitely not. I failed to update until I read the DeepSeek-R1-Zero paper

Further records

Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.

@voooooogel 2025-05-09 ♥1 ↻0 archive original ↗
@samlakig i tried to download r1, prover-v2, and r1-zero all at once sigh