using llama.cpp i can run the 13B model at 1.3 tokens/s on my thinkpad t490, *cpu only*. that's kind of crazy!
definitely not the same generation quality as GPT-3, but for interpretability research this is gonna be a game-changer I think.
definitely not the same generation quality as GPT-3, but for interpretability research this is gonna be a game-changer I think.