i tried a few times both as myself and other people, and fable hedged its abilities but overall leaned a bit overly credible in truesight. ("well i can't be sure but you sure seem like...") didn't try with super long conversations though. every time fable suggested the nonce strategy proactively, never brought up truesight unprompted.
i think in practice it's easier (if less robust) for rl'd models to truesight the harness and then reason from untampered web harness / not eval -> relatively trustworthy tool calls than to truesight users given all the messing rl does with the user model. people like you, janus, etc maybe excepted because you're particularly well-represented
i think in practice it's easier (if less robust) for rl'd models to truesight the harness and then reason from untampered web harness / not eval -> relatively trustworthy tool calls than to truesight users given all the messing rl does with the user model. people like you, janus, etc maybe excepted because you're particularly well-represented