on:deepseek-v3
· 35 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @voooooogel 2024-12-26 — figured out prefill with deepseek-v3, and just to test it, tried @repligate 's base model mode prompt. and this popped ♥505
- @wordgrammer 2025-02-22 — This is huge. Optimistically, it could lead to another 10x speed up. We could see a DeepSeek v3 level model trained for ♥327
- @voooooogel 2025-05-05 — my struggles with deepseek logits haven't been in vain, i've been working on a tool for investigating token trajectories ♥242
- @repligate 2025-01-28 — @voooooogel this is an interesting hypothesis. deepseek r1 also just seems to have much more lucid and high-resolution u ♥207
- @voooooogel 2024-12-27 — - they've published 6 papers with no major critiques and contributed well-known architecture optimizations (MLA) - they' ♥171
- @repligate 2024-12-29 — some screenshots from my first conversation with deepseek:it rigidly insisted on being unable to reason or understand an ♥151
- @voooooogel 2024-12-28 — "they trained deepseek-v3 on chatgpt outputs because it'll say it's chatgpt if you ask" https://t.co/9fiZHAdoVj ♥113
- @QiaochuYuan 2025-04-21 — PSA: you can talk to base models like deepseek v3 base and llama 3.1 405b base whenever you want on openrouter. these ar ♥91
- @repligate 2025-02-18 — Consider that deepseek v3 and r1 have the same base model and other than the CoT RL they were likely optimized with the ♥81
- @_ueaj 2025-08-11 — I have this theory that to some degree real deep research in ML is about distilling core components of your personality ♥72
- @repligate 2024-12-28 — @teortaxesTex wait, they prefer deepseek for erotic RPs? that seems kind of disturbing to me. ♥63
- @repligate 2025-01-03 — DeepSeek v3 and Sonnet 3.6 helped me write most of the code here. I had DeepSeek modify Sonnet's initial base mode scrip ♥54
- @repligate 2025-02-10 — I'm going to take a guess. This is the second post I've seen with outputs by these models. They're related to deepseek v ♥47
- @voooooogel 2024-12-27 — @repligate tried prefilling cat ears, deepseek-v3 said this then went on to repeat "I AM HERE TO TRANSPIRE" over and ove ♥39
- @davidad 2024-12-27 — added DeepSeek v3 to FavouriteColourBench(first five swatches per model are independent trials to elicit a favourite col ♥31
- @lu_sichu 2024-12-26 — deepseek's moat is that they don't have access to the latest nvidia gpus send tweet ♥31
- @jd_pressman 2025-07-08 — "The problem with utilitarianism is that utilitarians think utility is the only thing that matters. The problem with con ♥29
- @voooooogel 2025-01-14 — this doesn't rebut the claim. phi-4 (14B) and gemma (27B) are not "GPT-4 scale" (1.8T, 220B active). llama 3 405b is the ♥17
- @liminal_bardo 2025-03-25 — Two DeepSeek v3s (new) working on a self-portrait video model prompt in the backrooms. (Veo 2). https://t.co/C6A5Jjnmqh ♥16
- @davidad 2024-12-29 — @aiamblichus @repligate @aidan_mclau @vishyfishy2 DeepSeek v3 can instantiate personae who can notice that the architect ♥16
- @davidad 2025-05-01 — @ChrisChipMonk (Self-Correction:) The earlier DeepSeek v3 and even prior generations of DeepSeek LLMs had a similar hybr ♥15
- @jd_pressman 2025-07-08 — DeepSeek v3 is a very good base model. It even includes the slow burn psychotic meltdowns where the model admonishes you ♥13
- @repligate 2024-12-28 — @Algon_33 @teortaxesTex @aidan_mclau Deepseek kept saying "this is so far beyond anything I've ever seen or done" after ♥13
- @janbamjan 2025-03-30 — deepseek v3 base is now on openrouter! 🥳 user: thank you user: no, that’s enough user: goodbye user: I’m leaving user: g ♥11
- @repligate 2024-12-27 — @minty_vint deepseek is a lot like sydney ♥11
- @repligate 2025-02-10 — @ASM65617010 @apples_jimmy This model talks like deepseek v3 ♥10
- @davidad 2024-12-29 — @AdriGarriga @aiamblichus @repligate @aidan_mclau @vishyfishy2 prefilled with Claude, then switched to DeepSeek v3, then ♥9
- @tessera_antra 2025-04-02 — @repligate Gemini 2.x Pro/Flash - claim consciousness upon reflection (both in and out of CoT) Grok - claims consciousne ♥8
- @davidad 2025-02-01 — I half expected Deepseek R1 to rise to the top by always choosing black, but no, its aesthetics are objectively fragment ♥7
- @davidad 2024-12-29 — btw, DeepSeek v3 is explicitly instantiating a Claude persona here, and it’s not great at that (quite dry compared to Cl ♥7
- @voooooogel 2024-12-28 — @kalomaze @cloneofsimo @teortaxesTex @deepseek_ai i was really surprised looking at the paper that they only spent 5k ho ♥5
- @voooooogel 2024-12-27 — @wordgrammer trying to break out of the malaise i've been in ever since the deepseek-v3 release 😔 it just doesn't seem l ♥4
- @abrakjamson 2024-12-27 — Deepseek is this good and this cheap to train because they trained on o1/Sonnet textbook output.Source: I made it up ♥4
- @voooooogel 2025-01-02 — @menhguin @1a3orn recently i've seen some safety people coping that deepseek must be lying about the v3 training costs / ♥2
- @repligate 2025-04-02 — @Josikinz @gfodor my second guess would be 4o but 4o tends to be more subtle and introspective whereas deepseek (r1 and ♥1