@davidad 2026-03-20 ♥1 ↻0 original ↗
@JohnWittle My version of your hypothesis is that, since the training distribution clusters into tokenstreams which don’t look like evals (snippets of random books and web pages, natural chat conversations) and tokenstreams which do look like evals (coding problems, stilted chat
same thread: 1973032772127347039 1973033400090140992 1973054940995018917 1973056666074566875 1973064558387298467 1973064815766569252 1973073032655548579 1973375950101586310 1973377202705301532 1973377411464175906

author:davidad kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.