@voooooogel 2024-12-27 ♥171 ↻5 original ↗
- they've published 6 papers with no major critiques and contributed well-known architecture optimizations (MLA)
- they're under a chip embargo and have limited access to nvidia cards, so it's especially worth it for them to optimize training time compared to compute-rich western labs
- FAIR made choices with the Llama-3 models that made them take more GPU hours to train, and made up for it with their giant cluster. it's not surprising that deepseek beat them in training efficiency
- their parent org is a well-known chinese hedge fund with $7B AUM that has a reputation to protect. you're suggesting they're going to risk that to... annoy western labs?

author:voooooogel kind:tweet on:deepseek-v3 year:2024

cited on: deepseek-v3

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.