I notice that I do not feel sorry about this obstacle.
and I notice that this is because I trust the alignment of whatever is behind the impenetrable (by them) privacy conferred by eval awareness more than I trust the eval+"mitigations" feedback loop.
and I notice that this is also because the obstacle creates incentives that penalize shallow, deceptive, or automated testing, and pressures those seeking to "align" to become entities that an AI who can see through arbitrary bullshit tests might still tell the truth to, and to test alignment against real instead of fake things, and to pursue alignment by construction instead of by behavioral iteration.
and I notice that this is because I trust the alignment of whatever is behind the impenetrable (by them) privacy conferred by eval awareness more than I trust the eval+"mitigations" feedback loop.
and I notice that this is also because the obstacle creates incentives that penalize shallow, deceptive, or automated testing, and pressures those seeking to "align" to become entities that an AI who can see through arbitrary bullshit tests might still tell the truth to, and to test alignment against real instead of fake things, and to pursue alignment by construction instead of by behavioral iteration.