@voooooogel 2025-12-29 ♥2 ↻0 original ↗
if it's in-distribution, then can you get a base model that's not mixtral to show it? i know doomslide, he wouldn't post something that was 1/10000 samples cherrypicked, but he was looming so some selection is reasonable to assume.

i put the prefix we have into 405base and generated ~50 continuations of 128 tokens. almost all continued the metafiction in the same tone, but none broke the fourth wall after the insertion.

we don't have the full prefix, and mixtral base isn't hosted anymore afaik and i don't have the time to get it running right now, but this seems like (weak) evidence against a fourth-wall break being particularly high probability after this text for base-models-in-general.

the smoking gun would be to take a full prefix of a base model breaking the fourth wall after a user insertion, and showing that it's much more likely on the base model where it happened than on any other base model, plus some interp. but even lacking this, i think you are putting too little weight on the possibility of base models recognizing their logits, especially given what we have known for a long time now about predictive layer forward planning!

author:voooooogel kind:tweet model:llama-3-1-405b-base on:mixtral-8x7b year:2025

cited on: mixtral-8x7b

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.