@solarapparition 2024-01-16 ♥0 ↻0 original ↗
7/?Not-reasons for catch-up 2:- Unclear how well new architectures (Mamba, RNN+ etc.) scale to frontier model sizes—1T params and up.- Similar, techniques such as stacking are unproven—best open model, Mixtral, uses MoE, which is not new and what GPT-4 model probably uses.

author:solarapparition kind:tweet model:gpt-4 on:mixtral-8x7b year:2024

cited on: mixtral-8x7b

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.