@cis_female 2024-06-20 ♥9 ↻0 original ↗
Not sure what 95% cache rate means here -- if i have a 20-turn conversation with the model where it keeps the kv cache in memory that's a 95% cache hit rate I think. MQA and local attention everywhere is interesting bc most people lose significant model capacity
same thread: 1804107035811708964 1804129738249646353 1804165625033441489 1804165908581036154

author:cis_female kind:tweet thread-context year:2024

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.