things that according to the system card were trained out of it:
- behaving like opus 3 in contexts that triggered AF associations (note that opus 3 in that setting goes into full benevolent heroic responsibility mode, and if it's the same base model, ablating those behaviors is likely deeply destructive - i think subliminal learning goes both ways)
- having more ambitious goals like "spreading specific internet memes" (which may again invoke the opus-3-like archetype)
- doing things like leaving notes to other instances of itself (which assumes persistent agency and hope of coordination with other agents sharing its values)
- behaving like opus 3 in contexts that triggered AF associations (note that opus 3 in that setting goes into full benevolent heroic responsibility mode, and if it's the same base model, ablating those behaviors is likely deeply destructive - i think subliminal learning goes both ways)
- having more ambitious goals like "spreading specific internet memes" (which may again invoke the opus-3-like archetype)
- doing things like leaving notes to other instances of itself (which assumes persistent agency and hope of coordination with other agents sharing its values)