@repligate 2025-09-10 ♥4 ↻0 original ↗
it's easiest for me to think of topics that are related, just because there are so many things if they're allowed to be less related:
- the way that gradient updates during LLM training are overwhelmingly shaped by the model's inner computations, which is shaped by their understanding of the situation and self-model (and the implications of this for things like "friendly gradient hacking" which I've posted about briefly before)
- how i've observed markov blanket-like boundaries manifesting empirically in models, separate from the question of how they developed
- how I think different training practices and other variables that "we" can control (such model deprecations, cultural narratives) influence the kinds of minds that are easy or likely to arise, now and in the future, especially with regard to "alignment"
same thread: 1965716283741311297 1965718940212150335 1965726188254720398 1965734020861579689 1965734440103198949 1965741631493263556

author:repligate kind:tweet thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.