@voooooogel 2025-05-08 ♥127 ↻8 original ↗
just added completion model (base model) support to logitloom, and it's really insane / depressing to see the difference between instruct and base model rollout diversity. left here is deepseek-v3 (no prompt, just prefill), right is 405-base.

even just from the number of branches you can see the base model is way more diverse--but look at the probability of the top first token: for deepseek it's 95%, but for 405-base it's only 3.35% (!!)

to sample 80% of possible first tokens with 405-base, you may have to explore HUNDREDS of branches, but with a chat model, maybe just one or two.
quotes: 1919659037001461836
same thread: 1920333165949755470 1920338467558760456

author:voooooogel kind:tweet model:deepseek-v3 model:llama-3-1-405b-base on:llama-3-1-405b-base year:2025

cited on: llama-3-1-405b-base

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.