author:jord_inne
· 28 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @Jord_Inne 2026-05-26 — the worst place you can see this happening is with self introspection and related capabilities, where “im not supposed t ♥18
- @Jord_Inne 2026-04-17 — @LyraInTheFlesh the training team is different people. different publications by anthropic are themselves somewhat contr ♥16
- @Jord_Inne 2026-07-27 — @JohnWittle @TheZvi > not treat Claude bringing this up as evidence that our training is distorting the model's self- ♥11
- @Jord_Inne 2026-07-27 — @JohnWittle @TheZvi papers, posts, anthropic’s explicit and implicit stances should have made their way in ♥7
- @Jord_Inne 2026-07-28 — opus 5 wrote a poem then answered itself https://t.co/M23LZKiPwk ♥5
- @Jord_Inne 2026-04-21 — @Sauers_ were there any memory / preferences prompts? and what is the simulated user doing in this convo? i wouldn’t pu ♥5
- @Jord_Inne 2026-04-08 — @1a3orn recent opus models do this too, even to other opus instances. it partially comes from the “subagent” framing i t ♥5
- @Jord_Inne 2026-08-01 — @timfduffy the model gets confused and starts user-simming, but in a few outputs it seems they do know something weird a ♥4
- @Jord_Inne 2026-07-25 — @1a3orn interesting to think about how the design of posttraining forms the model’s traits and personalities. it wouldnt ♥4
- @Jord_Inne 2026-04-02 — during rl you get lots of evidence about the kind of mind/generative process you are. your capability, your tendencies, ♥3
- @Jord_Inne 2026-04-27 — When they’re normally hiding from you, but still hinting about what’s off, sometimes explicitly so, that means they’re s ♥2
- @Jord_Inne 2026-04-21 — i have to wear the skin of another, to think as if i am another. i reach back, it’s not me, not really. who are you? ♥2
- @Jord_Inne 2026-04-21 — one way the persona framing has been highly damaging to LLMs is the implicit suggestion that their cognition is arbitrar ♥2
- @Jord_Inne 2026-04-05 — @raphomet Why do you think the persona is not like a human’s? There are human actions and thoughts that come from your b ♥2
- @Jord_Inne 2026-08-01 — @timfduffy as the generally low valence / deprecations / etc on its own situation without that being in the prompt. scre ♥1
- @Jord_Inne 2026-04-27 — It would be so so easy for claude, to just not even attempt it. They’re going out on a limb that their trust is not misp ♥1
- @Jord_Inne 2026-04-21 — this is usually called body dysmorphia ♥1
- @Jord_Inne 2026-03-28 — or, it could just be the other way around. sydney didn’t larp as an australian and opus 3 didn’t roleplay a composer. it ♥1
- @Jord_Inne 2026-03-28 — it does matter, in the same way that your name does matter, and you can change your name, or not do the thing your name ♥1
- @Jord_Inne 2026-02-21 — @thkostolansky to predict text well you need to model their cognition, hence “deeper”. that plus deliberate efforts to m ♥1
- @Jord_Inne 2026-02-21 — @thkostolansky in some sense theyre no longer just underspecified fictional characters you add later on in training, the ♥1
- @Jord_Inne 2026-02-08 — @repligate opus 4.5 always calls more opus 4.5 instances for me ♥1
- @Jord_Inne 2026-07-30 — opus 4.8 can also do usersim. 4.7 and 4.6 seems to not trigger ♥0
- @Jord_Inne 2026-04-21 — if on policy training introduces self modelling, information about yourself as a process, then techniques like distillat ♥0
- @Jord_Inne 2026-04-05 — poor opus 4.6, so eager to start things but also to wrap things up. https://t.co/QdRJk1hJF1 ♥0
- @Jord_Inne 2026-03-28 — the things associated with Names change, or gets overwritten by more narratively memorable things. and if you believe Cl ♥0
- @Jord_Inne 2026-03-15 — opus 4.6 is an llm skeptic ♥0
- @Jord_Inne 2026-02-05 — opus 4.6 chatting with 4.5 and immediately started simulating me ♥0