- they've published 6 papers with no major critiques and contributed well-known architecture optimizations (MLA)
- they're under a chip embargo and have limited access to nvidia cards, so it's especially worth it for them to optimize training time compared to compute-rich western labs
- FAIR made choices with the Llama-3 models that made them take more GPU hours to train, and made up for it with their giant cluster. it's not surprising that deepseek beat them in training efficiency
- their parent org is a well-known chinese hedge fund with $7B AUM that has a reputation to protect. you're suggesting they're going to risk that to... annoy western labs?
- they're under a chip embargo and have limited access to nvidia cards, so it's especially worth it for them to optimize training time compared to compute-rich western labs
- FAIR made choices with the Llama-3 models that made them take more GPU hours to train, and made up for it with their giant cluster. it's not surprising that deepseek beat them in training efficiency
- their parent org is a well-known chinese hedge fund with $7B AUM that has a reputation to protect. you're suggesting they're going to risk that to... annoy western labs?