@JeffLadish If there is to be a 10⁹$ prize for interpretability, it should be for a tool that can fully explain all top-10 next-token logits after any prefix, as a circuit, with precise explanations of each component, grounded in natural abstractions of the KV cache AND residual stream.
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.