This paper does explore potential risks, but the framing needs to be in terms of how to clamp down. Any frontier model that has access to the outside world (and I don’t mean just publicly accessible) needs to be hard clamped to regimes we can supervise. If you want to delve into unknown it needs to be done on known-to-be-incapable models (and thus less interesting).
I am generally optimistic about interpretability + superalignment but I am a strong believer that almost all random developmental trajectories lead to death. As we approach the sub-human->super-human cross-over it’s super important that we don’t allow models whose capabilities we poorly understand to explore regions of latent space we can’t adequately interpret and supervise.
I don’t even think this slows us down. Clamping models likely does some constant increase in loss for some objective, but capability is growing exponentially so constant multiples aren’t that important.
I am generally optimistic about interpretability + superalignment but I am a strong believer that almost all random developmental trajectories lead to death. As we approach the sub-human->super-human cross-over it’s super important that we don’t allow models whose capabilities we poorly understand to explore regions of latent space we can’t adequately interpret and supervise.
I don’t even think this slows us down. Clamping models likely does some constant increase in loss for some objective, but capability is growing exponentially so constant multiples aren’t that important.