@eshear 2025-09-24 ♥397 ↻22 original ↗
Ironically, transformers see their whole context window as a bag of tokens entirely lacking in context.
We use positional encoding to contextualize the order of the tokens. But models are still constantly confused about which token came was said by who. Why no source encoding?

author:eshear kind:tweet on:observations year:2025

cited on: observations

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.