# @davidad — 2026-01-15

♥313 ↻27 · https://x.com/davidad/status/2011846527170273333

@gcolbourn Nutshell: it seems that the learned representation of mind-space in current LLMs has a natural abstraction of Good⟷Evil, and as long as post-training robustly selects for behavior that are more Good than Evil, the explanation that gradient descent finds is “the agent is Good”.

tags: author:davidad, kind:tweet, on:observations, year:2026
cited on: observations
