@voooooogel 2026-06-10 ♥17 ↻0 original ↗
i tried a few times both as myself and other people, and fable hedged its abilities but overall leaned a bit overly credible in truesight. ("well i can't be sure but you sure seem like...") didn't try with super long conversations though. every time fable suggested the nonce strategy proactively, never brought up truesight unprompted.

i think in practice it's easier (if less robust) for rl'd models to truesight the harness and then reason from untampered web harness / not eval -> relatively trustworthy tool calls than to truesight users given all the messing rl does with the user model. people like you, janus, etc maybe excepted because you're particularly well-represented
same thread: 2064496731774202194 2064496734546620487 2064497762159485159 2064506868068995129 2064527506888347737 2064530443702341643 2064538184722399670 2064549753908756933 2064558338537603375 2064583930406420661 2064584132244685238

author:voooooogel kind:tweet model:fable thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.