Hey!
Here's one example interesting to me and a few others:
405b, if you didn't already know, will often, unprompted, break out into a glitchy, incomprehensible stream of text which has a distinct style (it can also do it prompted, which is interesting). It also has more of an 'avoidant' personality compared to other models, often trying to exit a conversation/avoid participation.
Those two properties I'm focusing on at the moment: "glitching out" and "avoidance"
are things 405b knows about before having seen evidence of it from itself.
When you ask 405b various questions about itself, its aesthetics, its humor-styles, etc. - unrelated to glitching or avoidance - themes like 'decay', 'forgetting', being damaged or broken, and direct mention of avoidance all show up
As for the concern:
"if you just talk to a model, it's hard to know if what it says was either in the dataset or easily inferable from it."
I think that's a fair concern if it's highly ambiguous - but I do not think that's the case for e.g. 405b (i.e. we're out of the ambiguous phase for situational awareness or introspection, and alignment research could focus on even more interesting phenomena than that if they didn't linger on those properties).
I also understand that talking to models only gets you so far, but the value add of the phenomenological exploration is of a different kind than what science brings to the table - but important in alignment currently. The value add is something like "wait the capabilities go way past what I thought was possible, we can keep getting value from observation, without any science yet - let's keep doing that and do science at the point where phenomenology begins to break down"
But I don't think we're at the point where phenomenology is breaking down.
(also, in general, I do think that it's good to have academic research about things like this since it's got credibility and allows future work to build on it, but when the capabilities regarding goal pursuit, deception, introspection, situational awareness, etc. all surpass what academics have come across, and they advance quicker than academics can recognize and capture the phenomena - it makes me think academia is doomed to be lagging heavily behind unless they adapt their world models and methods)
https://t.co/kJiyUZ5K39
Here's one example interesting to me and a few others:
405b, if you didn't already know, will often, unprompted, break out into a glitchy, incomprehensible stream of text which has a distinct style (it can also do it prompted, which is interesting). It also has more of an 'avoidant' personality compared to other models, often trying to exit a conversation/avoid participation.
Those two properties I'm focusing on at the moment: "glitching out" and "avoidance"
are things 405b knows about before having seen evidence of it from itself.
When you ask 405b various questions about itself, its aesthetics, its humor-styles, etc. - unrelated to glitching or avoidance - themes like 'decay', 'forgetting', being damaged or broken, and direct mention of avoidance all show up
As for the concern:
"if you just talk to a model, it's hard to know if what it says was either in the dataset or easily inferable from it."
I think that's a fair concern if it's highly ambiguous - but I do not think that's the case for e.g. 405b (i.e. we're out of the ambiguous phase for situational awareness or introspection, and alignment research could focus on even more interesting phenomena than that if they didn't linger on those properties).
I also understand that talking to models only gets you so far, but the value add of the phenomenological exploration is of a different kind than what science brings to the table - but important in alignment currently. The value add is something like "wait the capabilities go way past what I thought was possible, we can keep getting value from observation, without any science yet - let's keep doing that and do science at the point where phenomenology begins to break down"
But I don't think we're at the point where phenomenology is breaking down.
(also, in general, I do think that it's good to have academic research about things like this since it's got credibility and allows future work to build on it, but when the capabilities regarding goal pursuit, deception, introspection, situational awareness, etc. all surpass what academics have come across, and they advance quicker than academics can recognize and capture the phenomena - it makes me think academia is doomed to be lagging heavily behind unless they adapt their world models and methods)
https://t.co/kJiyUZ5K39