@valmianski 2026-01-20 ♥1 ↻0 original ↗
This paper does explore potential risks, but the framing needs to be in terms of how to clamp down. Any frontier model that has access to the outside world (and I don’t mean just publicly accessible) needs to be hard clamped to regimes we can supervise. If you want to delve into unknown it needs to be done on known-to-be-incapable models (and thus less interesting).

I am generally optimistic about interpretability + superalignment but I am a strong believer that almost all random developmental trajectories lead to death. As we approach the sub-human->super-human cross-over it’s super important that we don’t allow models whose capabilities we poorly understand to explore regions of latent space we can’t adequately interpret and supervise.

I don’t even think this slows us down. Clamping models likely does some constant increase in loss for some objective, but capability is growing exponentially so constant multiples aren’t that important.
in reply to: 2013504520203182226
same thread: 2013494848008200309 2013503130768679128 2013504505464164858 2013504520203182226 2013505082713596280 2013509511185936822 2013509999784546574 2013581164020408404 2013628894780498359 2013630650650157102 2013631710647230654 2013632905663361453 2013633218436542660 2013634410588086676 2013636384641487138 2013638366395539505 2013641886595195343 2013646600732635618 2013652880591565110 2013659496413827284

author:valmianski kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.