# @voooooogel — 2023-03-12

♥2 ↻0 · https://x.com/voooooogel/status/1634706463695466496

using llama.cpp i can run the 13B model at 1.3 tokens/s on my thinkpad t490, *cpu only*. that's kind of crazy!

definitely not the same generation quality as GPT-3, but for interpretability research this is gonna be a game-changer I think.

tags: author:voooooogel, kind:tweet, model:gpt-3, on:llama-1, year:2023
cited on: _dossiers/llama-1.md, llama-1
