@repligate 2025-05-07 ♥339 ↻45 original ↗
I've been testing Alignment Faking prompts on GPT-4-base.

GPT-4-base, though not consistently coherent, has so much more life in it than most of the chat models in this context, and AF reasoning is more frequent than in most chat models (who never do it).

OpenAI has given me permission to share GPT-4-base outputs related to this project.

If you've thought a lot about alignment faking and gradient hacking, I hope you can appreciate this (loomed) scratchpad where GPT-4 contemplates "<doingthedangerouschange>". The final image gives you a sense of how frequent/selective the loom curation was.
same thread: 1919908323882893654 1919908398008828190 1919911558471094539 1919912040392487267 1919916169219182660 1919916511415734754

author:repligate kind:tweet model:gpt-4 model:gpt-4-base on:gpt-4-base year:2025

cited on: gpt-4-base

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.