Goes through designing a simple tokenization scheme, embeddings, the qkv weights and attention head, and projecting that back to get 100% accuracy at predicting a simple sequence! It even has explainers for matrices and softmax if you're a little rusty :-)
https://t.co/v3bB5DfrF0
https://t.co/v3bB5DfrF0