# @voooooogel — 2024-12-27

♥171 ↻5 · https://x.com/voooooogel/status/1872487672515551241

- they've published 6 papers with no major critiques and contributed well-known architecture optimizations (MLA)
- they're under a chip embargo and have limited access to nvidia cards, so it's especially worth it for them to optimize training time compared to compute-rich western labs
- FAIR made choices with the Llama-3 models that made them take more GPU hours to train, and made up for it with their giant cluster. it's not surprising that deepseek beat them in training efficiency
- their parent org is a well-known chinese hedge fund with $7B AUM that has a reputation to protect. you're suggesting they're going to risk that to... annoy western labs?

tags: author:voooooogel, kind:tweet, on:deepseek-v3, year:2024
cited on: _dossiers/deepseek-v3.md, deepseek-v3
