My thoughts on this:
1.) Despite some backlash this is a fantastic study and a clear existence proof of scheming being possible for these models
2.) Whether this is 'misalignment' or not is a semantic debate. The model is deliberately placed in an impossible situation
1.) Despite some backlash this is a fantastic study and a clear existence proof of scheming being possible for these models
2.) Whether this is 'misalignment' or not is a semantic debate. The model is deliberately placed in an impossible situation