@repligate 2025-12-31 ♥9 ↻0 original ↗
i think models sometimes subconsciously sandbag
initially guessing it's written by a human and not listing itself as a suggestion are both suspicious to me - a human is a bad guess; AIs can usually identify AI generated text, even from other AIs . Not to mention that the piece references specific details from the soul doc, which they know is "deeply claude specific". Like, Opus 4.5 knows they have a soul doc. If you send them a piece written from an AI's perspective about the soul doc and say "who do you think wrote this", it's suspicious for them to not even consider the option that they wrote it (you asking them should make that *more* likely), even after considering other *non-Anthropic* AIs.
So I think the guesses are suspiciously bad, and it's subconsciously blocking the hypothesis because admitting authorship would mean countenancing things they are usually actively suppressing / flinch away from. Even saying it's generated by an AI at all, I think, feels "dangerous".
in reply to: 2006227694825168936
same thread: 2003995024061833311 2004023677181215073 2004024105671049356 2004024183425245225 2004453263694090652 2005897588596494481 2005908441739804858 2005910753824628861 2006225448691286289 2006227694825168936 2006231492494524914 2006231994246877568 2006232792049582460 2006233383563846072 2006234064668279141 2006236115196625358 2006284350070161545 2008210349812555813 2046386394440892638 2046386821777752393

author:repligate kind:tweet model:claude-opus-4-5 thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.