@voooooogel 2023-03-12 ♥2 ↻0 original ↗
using llama.cpp i can run the 13B model at 1.3 tokens/s on my thinkpad t490, *cpu only*. that's kind of crazy!

definitely not the same generation quality as GPT-3, but for interpretability research this is gonna be a game-changer I think.

author:voooooogel kind:tweet model:gpt-3 on:llama-1 year:2023

cited on: llama-1

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.