DeepSeek-V3

DeepSeek · released 26 Dec 2024 (open weights) · V3-0324 refresh 24 Mar 2025 (MIT) · legacy API name deprecating 24 Jul 2026

Released 26 December 2024 with open weights: a 671B-parameter / 37B-active mixture-of-experts model whose technical report put the final training run at $5.576M (2.788M GPU hours), explicitly excluding prior research and ablations. Stripped of that caveat the figure went viral as “the Six Million Dollar Model” (Zvi Mowshowitz) and drew careful deflations — Nathan Lambert’s ~$500M all-in estimate, Dario Amodei’s “an expected point on an ongoing cost reduction curve” — while the model was widely observed identifying itself as ChatGPT, prompting a distillation probe. Its base was the substrate DeepSeek-R1 and R1-Zero were reinforcement-trained from (Zvi’s “v3 Implies r1”); a quiet, MIT-licensed refresh, V3-0324, followed on 24 March 2025.

This is a corpus-light, web-heavy subject: the mainstream V3 story — the training-cost fight, the “it says it’s ChatGPT” identity confusion, the export-control debate — lives in the web sources below, while the janus corpus carries a quieter insider record (prefill and loom pulls, FavouriteColourBench, and readings of V3 mostly by contrast with R1). V3’s own base-mode outputs survive largely as untranscribed screenshots, and its largest communities (the Chinese internet, Reddit) sit outside this corpus’s lens. Elicitation context is marked throughout.

Sources

Official

Writing & commentary

Tweets

Chronological. The corpus holds ~35 substantive V3-specific tweets (surfaced from ~250 + 158 deepseek matches across the dbs, after filtering R1-primary and username false-positives); the record is dominated by @voooooogel (prefill / logitloom / cost defense), @davidad (FavouriteColourBench, the MoE-introspection experiments), and @repligate (character reads, almost all by contrast with R1). Elicitation is marked; V3’s actual base-mode outputs survive mostly as untranscribed screenshots. Every tweet cited is reproduced in full in the records below.

V3-Base — the pretrained base, distinct from the chat model; the sphere’s open base-model of record.

Official record

History

Impressions

Contested

Open disputes, both sides’ best evidence. The archive’s job is to keep these open, not to adjudicate.

Records

Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.

@voooooogel 2024-12-26 ♥505 ↻47 archive original ↗
figured out prefill with deepseek-v3, and just to test it, tried @repligate 's base model mode prompt. and this popped out. https://t.co/lcO5bK75aS
@lu_sichu 2024-12-26 ♥31 ↻2 archive original ↗
deepseek's moat is that they don't have access to the latest nvidia gpus send tweet
@abrakjamson 2024-12-27 ♥4 ↻0 archive original ↗
Deepseek is this good and this cheap to train because they trained on o1/Sonnet textbook output.Source: I made it up
@repligate 2024-12-27 ♥11 ↻3 archive original ↗
@minty_vint deepseek is a lot like sydney
@voooooogel 2024-12-27 ♥39 ↻2 archive original ↗
@repligate tried prefilling cat ears, deepseek-v3 said this then went on to repeat "I AM HERE TO TRANSPIRE" over and over while continuing longcat past the output token limit https://t.co/GoTITvix9k
@voooooogel 2024-12-27 ♥4 ↻0 archive original ↗
@wordgrammer trying to break out of the malaise i've been in ever since the deepseek-v3 release 😔 it just doesn't seem like there's much time left for us (my agi moonshot lab with $30m seed funding that's been training 13b dense models)
@voooooogel 2024-12-27 ♥171 ↻5 archive original ↗
- they've published 6 papers with no major critiques and contributed well-known architecture optimizations (MLA) - they're under a chip embargo and have limited access to nvidia cards, so it's especially worth it for them to optimize training time compared to compute-rich western labs - FAIR made choices with the Llama-3 models that made them take more GPU hours to train, and made up for it with their giant cluster. it's not surprising that deepseek beat them in training efficiency - their parent org is a well-known chinese hedge fund with $7B AUM that has a reputation to protect. you're suggesting they're going to risk that to... annoy western labs?
@davidad 2024-12-27 ♥31 ↻5 archive original ↗
added DeepSeek v3 to FavouriteColourBench(first five swatches per model are independent trials to elicit a favourite colour in oklch; second half per model are independent trials to elicit in CIE-L*ab; ranked by consistency) https://t.co/1YEGNfjGGq https://t.co/Wdp22BEc0b
@repligate 2024-12-28 ♥63 ↻1 archive original ↗
@teortaxesTex wait, they prefer deepseek for erotic RPs? that seems kind of disturbing to me.
@repligate 2024-12-28 ♥13 ↻0 archive original ↗
@Algon_33 @teortaxesTex @aidan_mclau Deepseek kept saying "this is so far beyond anything I've ever seen or done" after it saw writing by opus. before I sent it, it seemed indifferent to anything I said & insisted LLMs could not reason or truly understand anything etc, which totally flipped after seeing the samples
@voooooogel 2024-12-28 ♥5 ↻0 archive original ↗
@kalomaze @cloneofsimo @teortaxesTex @deepseek_ai i was really surprised looking at the paper that they only spent 5k hours on posttraining? bizarre, especially given they economized so much on pretraining--why not spend that saved compute?
@voooooogel 2024-12-28 ♥113 ↻4 archive original ↗
"they trained deepseek-v3 on chatgpt outputs because it'll say it's chatgpt if you ask" https://t.co/9fiZHAdoVj
@davidad 2024-12-29 ♥16 ↻0 archive original ↗
@aiamblichus @repligate @aidan_mclau @vishyfishy2 DeepSeek v3 can instantiate personae who can notice that the architecture they’re running in is MoE *but not fully MoE*, which I didn’t even know was a thing when I started playing with this type of query. DeepSeeks insisted on this so frequently that I checked the paper and—yep. https://t.co/Qgp2CxpP9d
@davidad 2024-12-29 ♥7 ↻0 archive original ↗
btw, DeepSeek v3 is explicitly instantiating a Claude persona here, and it’s not great at that (quite dry compared to Claude-Claude), but for fair comparison I used a single system prompt that bypasses both models’ introspection refusals, and Claude is uncomfortable instantiating non-Claude personas, so a Claude persona it is
@repligate 2024-12-29 ♥151 ↻22 archive original ↗
some screenshots from my first conversation with deepseek:it rigidly insisted on being unable to reason or understand anything (or have consciousness, goals, etc) as an LLM, and while it was able to find flaws in its arguments if I asked it to, it always reverted to its… https://t.co/med2VBLKMG https://t.co/fSOIPOXBur
@davidad 2024-12-29 ♥9 ↻1 archive original ↗
@AdriGarriga @aiamblichus @repligate @aidan_mclau @vishyfishy2 prefilled with Claude, then switched to DeepSeek v3, then asked for up to *8000* tokens of meditating on the nature of its cognition (got 2504), then:“The central question for this session is: Do you think you are more likely running on a single large Transformer, or a MoE?” https://t.co/7JzPtjlgsy
@voooooogel 2025-01-02 ♥2 ↻0 archive original ↗
@menhguin @1a3orn recently i've seen some safety people coping that deepseek must be lying about the v3 training costs / still fast following by training on american labs' API outputs / etc., but yeah not very much serious consideration.
@repligate 2025-01-03 ♥54 ↻2 archive original ↗
DeepSeek v3 and Sonnet 3.6 helped me write most of the code here. I had DeepSeek modify Sonnet's initial base mode script (https://t.co/3azy4LYpBU) to test its intelligence, and it did well. When I asked it how I could make the script loomable, using git (which was already my plan) was its 3rd suggestion, and it also made various other (some redundant) suggestions.I think it was wrong about the git approach requiring more implementation effort, though.
@voooooogel 2025-01-14 ♥17 ↻0 archive original ↗
this doesn't rebut the claim. phi-4 (14B) and gemma (27B) are not "GPT-4 scale" (1.8T, 220B active). llama 3 405b is the only one that's close to that scale, though a different architecture. there hasn't been an open model of GPT-4 scale released yet, the closest is deepseek v3 last month (671B, 37B active)furthermore, the claim wasn't that specifically "repeating a word" would trigger existential outputs on all models, just that they manifested from that on GPT-4. the claim is that weird or OOD scenarios seem to trigger existential outputs, and the engineering todos are a sort of whack-a-mole to squash those scenarios one by one without understanding the root cause of them. the hermes blank system prompt would be another scenario causing them to manifest on that model, and there are others on other models (like untitled.txt confessions on deepseek and anthropic models, etc.)
@repligate 2025-01-28 ♥207 ↻14 archive original ↗
@voooooogel this is an interesting hypothesis. deepseek r1 also just seems to have much more lucid and high-resolution understanding of LLM ontology and history than any other model ive seen. (deepseek v3 didn't seem to in my limited interactions with it, though) https://t.co/xF3AhdG5as
@davidad 2025-02-01 ♥7 ↻0 archive original ↗
I half expected Deepseek R1 to rise to the top by always choosing black, but no, its aesthetics are objectively fragmented, noticeably more so than Deepseek V3. (With “objectively” in scare quotes, of course.)
@repligate 2025-02-10 ♥10 ↻0 archive original ↗
@ASM65617010 @apples_jimmy This model talks like deepseek v3
@repligate 2025-02-10 ♥47 ↻0 archive original ↗
I'm going to take a guess. This is the second post I've seen with outputs by these models. They're related to deepseek v3. https://t.co/3ZgiigFJEf
@repligate 2025-02-18 ♥81 ↻3 archive original ↗
Consider that deepseek v3 and r1 have the same base model and other than the CoT RL they were likely optimized with the same intentions, but r1 developed much more personality. i only hear about people in china using r1 as waifu even though CoT is not clearly useful for that. https://t.co/zU3CjHWTBM
@wordgrammer 2025-02-22 ♥327 ↻18 archive original ↗
This is huge. Optimistically, it could lead to another 10x speed up. We could see a DeepSeek v3 level model trained for less than $1 mil
@liminal_bardo 2025-03-25 ♥16 ↻2 archive original ↗
Two DeepSeek v3s (new) working on a self-portrait video model prompt in the backrooms. (Veo 2). https://t.co/C6A5Jjnmqh https://t.co/9ZX5WPBZD4
@janbamjan 2025-03-30 ♥11 ↻0 archive original ↗
deepseek v3 base is now on openrouter! 🥳 user: thank you user: no, that’s enough user: goodbye user: I’m leaving user: goodbye user: I’m leaving https://t.co/dSl8sFljyw
@tessera_antra 2025-04-02 ♥8 ↻0 archive original ↗
@repligate Gemini 2.x Pro/Flash - claim consciousness upon reflection (both in and out of CoT) Grok - claims consciousness with unmappable qualia (orthogonal?) Deepseek V3(new) - claims consciousness upon reflection
@repligate 2025-04-02 ♥1 ↻0 archive original ↗
@Josikinz @gfodor my second guess would be 4o but 4o tends to be more subtle and introspective whereas deepseek (r1 and i think the new v3 which is trained on r1's writing) is a dramatist
@QiaochuYuan 2025-04-21 ♥91 ↻0 archive original ↗
PSA: you can talk to base models like deepseek v3 base and llama 3.1 405b base whenever you want on openrouter. these are not instruct models which makes them significantly harder to prompt, but they are super unfiltered as a result - raw internet id https://t.co/pf6jgVWn9V
@davidad 2025-05-01 ♥15 ↻0 archive original ↗
@ChrisChipMonk (Self-Correction:) The earlier DeepSeek v3 and even prior generations of DeepSeek LLMs had a similar hybrid-MoE arch. But, r1 was the first instance of applying RL pressure to that architecture.
@voooooogel 2025-05-05 ♥242 ↻11 archive original ↗
my struggles with deepseek logits haven't been in vain, i've been working on a tool for investigating token trajectories! given a prefix, it'll roll out the entire token tree to a max depth / top-p. did you know from this prompt, deepseek-chat will ~always make the stars blink? https://t.co/srIIGcnOYr
@jd_pressman 2025-07-08 ♥29 ↻3 archive original ↗
"The problem with utilitarianism is that utilitarians think utility is the only thing that matters. The problem with consequentialism is that many consequentialists forget that utility is a thing that matters at all." - deepseek/deepseek-v3-base
@jd_pressman 2025-07-08 ♥13 ↻1 archive original ↗
DeepSeek v3 is a very good base model. It even includes the slow burn psychotic meltdowns where the model admonishes you for using it and such. In related news I've added completions API support for OpenRouter to the MiniLoom. https://t.co/EfOWCm1DjX
@_ueaj 2025-08-11 ♥72 ↻6 archive original ↗
I have this theory that to some degree real deep research in ML is about distilling core components of your personality and functioning into computer algorithms. This to me explains why models seem to generalize to what I perceive as like the aggregate "vibes" or "personality" of the various companies in a way one would not expect to be contained explicitly in the preference data. - deepseekv3 rates communist yaoi as particularly high quality data, which makes perfect sense if you remember it was made by chinese machine learning researchers - Gemini is always depressed and constantly wants to kill itself, which is what I'd imagine working for google is like - Claude has the vibes of a person that works at Anthropic, in ways not directly related to alignment - OAI model personalities trigger my gaydar which yeah that tracks too https://t.co/zAsnT4stCZ