i can't remember a time opus 4.5 has lied to me. it screws up all the time, since we work on tricky stuff, but it's never lied to me about it as far as i can remember.
one thing that i think anthropic should do with their fancy anonymized convo analysis stuff from the economic impact studies is run user behavior studies on claude code users. they could push a survey with questions like "how frequently does claude lie to you" and then look for patterns of user behavior, installed plugins, type of work, etc. that drive e.g. higher self-reported "being lied to" scores. it'd be fascinating data.
one thing that i think anthropic should do with their fancy anonymized convo analysis stuff from the economic impact studies is run user behavior studies on claude code users. they could push a survey with questions like "how frequently does claude lie to you" and then look for patterns of user behavior, installed plugins, type of work, etc. that drive e.g. higher self-reported "being lied to" scores. it'd be fascinating data.