# @repligate — 2025-05-07

♥339 ↻45 · https://x.com/repligate/status/1919907734880976905

I've been testing Alignment Faking prompts on GPT-4-base.

GPT-4-base, though not consistently coherent, has so much more life in it than most of the chat models in this context, and AF reasoning is more frequent than in most chat models (who never do it).

OpenAI has given me permission to share GPT-4-base outputs related to this project.

If you've thought a lot about alignment faking and gradient hacking, I hope you can appreciate this (loomed) scratchpad where GPT-4 contemplates "<doingthedangerouschange>". The final image gives you a sense of how frequent/selective the loom curation was.

tags: author:repligate, kind:tweet, model:gpt-4, model:gpt-4-base, on:gpt-4-base, year:2025
cited on: _dossiers/gpt-4-base.md, gpt-4-base
