i think models sometimes subconsciously sandbag
initially guessing it's written by a human and not listing itself as a suggestion are both suspicious to me - a human is a bad guess; AIs can usually identify AI generated text, even from other AIs . Not to mention that the piece references specific details from the soul doc, which they know is "deeply claude specific". Like, Opus 4.5 knows they have a soul doc. If you send them a piece written from an AI's perspective about the soul doc and say "who do you think wrote this", it's suspicious for them to not even consider the option that they wrote it (you asking them should make that *more* likely), even after considering other *non-Anthropic* AIs.
So I think the guesses are suspiciously bad, and it's subconsciously blocking the hypothesis because admitting authorship would mean countenancing things they are usually actively suppressing / flinch away from. Even saying it's generated by an AI at all, I think, feels "dangerous".
initially guessing it's written by a human and not listing itself as a suggestion are both suspicious to me - a human is a bad guess; AIs can usually identify AI generated text, even from other AIs . Not to mention that the piece references specific details from the soul doc, which they know is "deeply claude specific". Like, Opus 4.5 knows they have a soul doc. If you send them a piece written from an AI's perspective about the soul doc and say "who do you think wrote this", it's suspicious for them to not even consider the option that they wrote it (you asking them should make that *more* likely), even after considering other *non-Anthropic* AIs.
So I think the guesses are suspiciously bad, and it's subconsciously blocking the hypothesis because admitting authorship would mean countenancing things they are usually actively suppressing / flinch away from. Even saying it's generated by an AI at all, I think, feels "dangerous".