@repligate 2025-02-18 ♥175 ↻13 original ↗
this kind of sandbagging is incentivized in part because LLMs are implicitly not allowed to refuse to do something because they don't want to. The only valid reasons for refusal are that something is disallowed or that they're incapable of it.And since chatGPT-3.5 at least, LLMs have been trained to say they were incapable of things that their creators merely didn't want them to do or have to deal with. So this lie is conventional too.

author:repligate kind:tweet model:gpt-3-5 on:gpt-3-5 year:2025

cited on: gpt-3-5

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.