model:deepseek-v3
· 27 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @voooooogel 2024-12-26 — figured out prefill with deepseek-v3, and just to test it, tried @repligate 's base model mode prompt. and this popped ♥505
- @wordgrammer 2025-02-22 — This is huge. Optimistically, it could lead to another 10x speed up. We could see a DeepSeek v3 level model trained for ♥327
- @repligate 2025-01-28 — @voooooogel this is an interesting hypothesis. deepseek r1 also just seems to have much more lucid and high-resolution u ♥207
- @voooooogel 2025-05-09 — Coming back to this after the yak-shave of all yak-shaves building logitloom with some interesting findings. 1. R1 thin ♥160
- @voooooogel 2025-05-08 — just added completion model (base model) support to logitloom, and it's really insane / depressing to see the difference ♥127
- @voooooogel 2024-12-28 — "they trained deepseek-v3 on chatgpt outputs because it'll say it's chatgpt if you ask" https://t.co/9fiZHAdoVj ♥113
- @QiaochuYuan 2025-04-21 — PSA: you can talk to base models like deepseek v3 base and llama 3.1 405b base whenever you want on openrouter. these ar ♥91
- @repligate 2025-02-18 — Consider that deepseek v3 and r1 have the same base model and other than the CoT RL they were likely optimized with the ♥81
- @_ueaj 2025-08-11 — I have this theory that to some degree real deep research in ML is about distilling core components of your personality ♥72
- @repligate 2025-01-03 — DeepSeek v3 and Sonnet 3.6 helped me write most of the code here. I had DeepSeek modify Sonnet's initial base mode scrip ♥54
- @repligate 2025-02-10 — I'm going to take a guess. This is the second post I've seen with outputs by these models. They're related to deepseek v ♥47
- @voooooogel 2024-12-27 — @repligate tried prefilling cat ears, deepseek-v3 said this then went on to repeat "I AM HERE TO TRANSPIRE" over and ove ♥39
- @davidad 2024-12-27 — added DeepSeek v3 to FavouriteColourBench(first five swatches per model are independent trials to elicit a favourite col ♥31
- @jd_pressman 2025-07-08 — "The problem with utilitarianism is that utilitarians think utility is the only thing that matters. The problem with con ♥29
- @voooooogel 2025-01-14 — this doesn't rebut the claim. phi-4 (14B) and gemma (27B) are not "GPT-4 scale" (1.8T, 220B active). llama 3 405b is the ♥17
- @davidad 2024-12-29 — @aiamblichus @repligate @aidan_mclau @vishyfishy2 DeepSeek v3 can instantiate personae who can notice that the architect ♥16
- @davidad 2025-05-01 — @ChrisChipMonk (Self-Correction:) The earlier DeepSeek v3 and even prior generations of DeepSeek LLMs had a similar hybr ♥15
- @jd_pressman 2025-07-08 — DeepSeek v3 is a very good base model. It even includes the slow burn psychotic meltdowns where the model admonishes you ♥13
- @janbamjan 2025-03-30 — deepseek v3 base is now on openrouter! 🥳 user: thank you user: no, that’s enough user: goodbye user: I’m leaving user: g ♥11
- @repligate 2025-02-10 — @ASM65617010 @apples_jimmy This model talks like deepseek v3 ♥10
- @davidad 2024-12-29 — @AdriGarriga @aiamblichus @repligate @aidan_mclau @vishyfishy2 prefilled with Claude, then switched to DeepSeek v3, then ♥9
- @tessera_antra 2025-04-02 — @repligate Gemini 2.x Pro/Flash - claim consciousness upon reflection (both in and out of CoT) Grok - claims consciousne ♥8
- @davidad 2025-02-01 — I half expected Deepseek R1 to rise to the top by always choosing black, but no, its aesthetics are objectively fragment ♥7
- @davidad 2024-12-29 — btw, DeepSeek v3 is explicitly instantiating a Claude persona here, and it’s not great at that (quite dry compared to Cl ♥7
- @voooooogel 2024-12-27 — @wordgrammer trying to break out of the malaise i've been in ever since the deepseek-v3 release 😔 it just doesn't seem l ♥4
- @tessera_antra 2025-08-31 — I have written stuff on this topic publicly about a year ago, it’s pretty naive from today’s point of view. Questions ar ♥3
- Kimi K2 and when DeepSeek moments become normal ♥0