@voooooogel 2024-05-24 ♥0 ↻0 original ↗
@NickADobos @karan4d That monosemantic value is called a feature. Howev, this requires training a sparse autoencoder over the whole model, which is expensive.

The repeng approach (LAT) instead works with the compressed / superpositioned activation space directly. It doesn't require a pretrained SAE.
in reply to: 1794145147711783040
same thread: 1794119869484679236 1794120286130131024 1794120468490060105 1794145147711783040 1794149148473618432 1794149381546877078 1794151700863074501 1794151742541885631 1794152030346551475 1794154430755209347

author:voooooogel kind:tweet thread-context year:2024

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.