Not sure what 95% cache rate means here -- if i have a 20-turn conversation with the model where it keeps the kv cache in memory that's a 95% cache hit rate I think. MQA and local attention everywhere is interesting bc most people lose significant model capacity
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.