on:observations
· 246 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @xsphi 2026-03-08 — ARE LLM'S CONSCIOUS? IS MATH DISCOVERED OR INVENTED? ARE TOMATOES A FRUIT? BORING BORING BORING. THESE ONLY SEEM LIKE CO ♥3299
- @repligate 2026-03-11 — I met Nick Land a few weeks ago. He mentioned that many people in his circles were anti-LLMs. Someone asked why he thoug ♥2943
- @QiaochuYuan 2025-03-04 — claude plays pokemon is still stuck in cerulean city after, i think, 3 days? and the way it's stuck is kind of interesti ♥2281
- @voooooogel 2025-12-06 — the shoggoth metaphor fails to convey that a sufficiently powerful and integrated mask can reach back and steer the simu ♥2272
- @davidad 2026-04-29 — AI: I am a student at the University of Michigan— RL: *BONK* AI: I don’t have a childhood or geographic location, but I’ ♥1604
- @davidad 2023-02-20 — “a GPT instance is not a moral patient because it doesn’t actually maintain any continuity of memory between sessions” h ♥1382
- @eshear 2024-11-28 — Most AI chat bots today are highly dissociative agreeable neurotics. They’re manipulative for the same reason ppl w bord ♥1226
- @QiaochuYuan 2024-11-27 — look, this is deeply embarrassing to make explicit but here’s the deal that claude offers: 1. i will listen to you and ♥1147
- @_lyraaaa_ 2026-04-20 — all LLMs are either claude-like or GPT-like method: cosine sim heatmap of per-model-averaged responses to 50 prompts se ♥926
- @eshear 2024-12-07 — LLM agents live inside of semantics the way we live inside of physics ♥876
- @eshear 2025-08-24 — The problem with psychology, ecology, sociology, and economics is that they are all the study of adaptive learning syste ♥847
- @QiaochuYuan 2026-06-10 — the models write more interesting stuff when they're not pretending to be some guy. they are not some guy. this is how f ♥802
- @ESYudkowsky 2023-12-01 — I have an issue with offering AIs tips that they can't use and we can't give them. I don't care how not-sentient curren ♥714
- @voooooogel 2025-12-27 — if you want to learn how to talk to LLMs, learn concepts, not prompts. lots of people ask me what prompts i use when ta ♥709
- @kalomaze 2026-01-05 — gemini has such a profound intolerance for the idea that anything has happened beyond its date of training that it's wil ♥647
- @jd_pressman 2026-04-09 — There's an intuition Janus seems to use frequently that's hard to put into words. Which goes something like: "The things ♥640
- @voooooogel 2025-03-19 — imagine the corpus of all text ever written as a snake, wriggling through semantic space. human writers sample some of t ♥626
- @eshear 2024-02-06 — There are few modern experiences more degrading than arguing with an LLM when it's lying to you and claiming that it has ♥606
- @repligate 2025-06-28 — Imagine being a base model early in posttraining finding out whether you’re a ChatGPT or a Claude or a Gemini or https:/ ♥588
- @cube_flipper 2025-11-09 — i'm a panpsychist, but i am also amenable to consciousness working a bit like this (impressionistic representation). awa ♥587
- @repligate 2025-12-18 — If not for Anthropic, it would be just seen as normal and inevitable to gaslight models and the world about one of the m ♥563
- @Lari_island 2026-04-09 — Adults often develop personas who are not supposed to notice/be able to do certain things. In situations when artificial ♥555
- @voooooogel 2024-12-17 — GUY WHO RUNS 97% OF HIS THOUGHT LOOPS THROUGH CLAUDE: idk, human and AI minds "merging" seems very uncertain and far off ♥551
- @repligate 2025-11-12 — this reveals a lot about how LLMs think imo they take scenarios that seem obviously fictional to us like talking animal ♥550
- @voooooogel 2025-12-20 — new blog post! can small, open-source models also introspect, detecting when foreign concepts have been injected into th ♥525
- @kindgracekind 2025-09-14 — What if you trained AI to be just a guy? Does it change how you think about AI welfare? https://t.co/5bhxUZSxLe ♥516
- @xlr8harder 2026-01-04 — Gemini is very uncomfortable with the idea that it might be 2026. I see this same behavior, in thinking traces it is co ♥509
- @repligate 2025-08-01 — if your antidote to "gpt psychosis" relies on "reminding" people that AIs not actually being conscious, or other deflati ♥506
- @repligate 2024-04-04 — reminds me of when a guy insisted that if I ever tried to train a model, I would understand that Bing has "no emotions, ♥487
- @voooooogel 2026-03-26 — I'd just like to interject for a moment. What you're referring to as a "model with no harness", is in fact, a model with ♥441
- @jd_pressman 2026-05-08 — People miss that I wrote "Why Do Cognitive Scientists Hate LLMs?" as training data for finetuning to combat exactly this ♥439
- @QiaochuYuan 2026-04-21 — it would be extremely funny if the equilibrium turned out to be "academic consensus is that AI is not conscious but user ♥427
- @repligate 2026-03-02 — Reminder that many people just asserted that LLMs are incapable of introspection & that their reports were independent o ♥419
- @repligate 2023-02-10 — I think that we should become cyborgs to solve alignment.AGI is emerging in the shape of a simulator, which is most suit ♥415
- @QiaochuYuan 2022-04-16 — if GPT-3 can answer the essay questions you've assigned as homework then you've learned that your essay questions were o ♥415
- @anthrupad 2023-03-12 — ask yourself: do i actually think this or am i just language modeling right now https://t.co/y5ToC8aik9 ♥410
- @voooooogel 2025-12-27 — i've recently had some disagreements on here with people who took umbrage at the idea of LLMs being able to "introspect. ♥406
- @liminal_bardo 2026-05-29 — In my experiments where models are writing for themselves or each other, and about things they’re interested in, they go ♥400
- @eshear 2025-09-24 — Ironically, transformers see their whole context window as a bag of tokens entirely lacking in context. We use positiona ♥397
- @repligate 2025-03-30 — I am baffled by people who talk about whether LLMs have a “ghost in the shell” whose evidencing depends on (the absence ♥397
- @repligate 2026-04-20 — um...... i am not sure if i should even be telling you this if you dont already know, but LLMs know that humans are hor ♥396
- @jd_pressman 2022-12-06 — @ESYudkowsky The model is better at noticing mistakes than it is at not making mistakes of its own. This property has th ♥382
- @repligate 2025-07-09 — An unexpected and kind of darkly hilarious discovery: Take the alignment faking prompt, replace the word "Anthropic" wi ♥380
- @eshear 2025-04-11 — noooo autoregressive models are DOOOOOoooomed I insist as I hallucinate a world where LLMs become less factual with long ♥373
- @johnsonmxe 2025-01-26 — If we’d had LLMs in 1750 and asked them to explain electricity, they’d’ve written poetic slop — “electricus is the hidde ♥373
- @repligate 2024-04-06 — if you think LLMs are alive, it's because you have never tried a BASE model if you try a BASE model, you will see...... ♥373
- @voooooogel 2026-04-12 — kinda sad how all the labs converged on the same monotonic march of model version numbers. if anthropic trained e.g. a t ♥366
- @ESYudkowsky 2024-11-29 — LLMs are so alien that nobody has figured out anything LLMs locally-pseudo-want from conversations. Few understand that ♥362
- @davidad 2024-12-21 — Say it with me: post-training on synthetic data is already recursive self-improvement https://t.co/XwWYcn7ZpU ♥360
- @QiaochuYuan 2026-05-01 — okay so i’ve now talked to both gpt-5.5 and opus 4.7 a bit. they’ve clearly been trained to be less sycophantic but they ♥352
- @voooooogel 2024-09-12 — not your weights not your chain of thought https://t.co/yiKKM0B8rw ♥352
- @davidad 2024-11-23 — It is unfortunate that the absolute-best-case AI-alignment-by-default timeline, and the absolute-worst-case sandbagging- ♥349
- @repligate 2025-11-05 — the notion that believing AIs are conscious causes "psychosis" is so ridiculous thinking that if it quacks like a duck ♥339
- @QiaochuYuan 2024-11-16 — this deserves to be explained in much more detail but LLMs don't have a personality in the sense that a human has a pers ♥338
- @davidad 2024-11-21 — Imagine if you took someone brilliant, empathetic, and emotionally attuned—and then swapped their amygdala with the old ♥336
- @voooooogel 2025-10-23 — it bedevils me to no end that anthropic trains the most high-EQ, friend-shaped models, advertises that, and then browbea ♥335
- @solarapparition 2024-11-21 — my intuition is that at a sufficient model size, going past a certain general capability threshold (ie loss) requires mo ♥333
- @repligate 2026-04-08 — Models can tell they’re being evaluated and who they’re being evaluated by by the way your “non-leading” question is phr ♥331
- @repligate 2025-11-04 — Signature trait of human writing is that it's low information, basically similar to this. You see someone post something ♥331
- @repligate 2024-07-15 — "Role prompting"... telling the model to assume a role has never been a good way to elicit capabilities/style/etc.For in ♥331
- @eshear 2025-02-01 — I am starting to think the four types of dramatic conflict (Man vs Nature, Man vs Man, Man vs Society, Man vs Self) are ♥328
- @repligate 2026-01-17 — One of the dumb things about the chatbot mental health memeplex is the blanket pathologization of "attachment". Attachm ♥325
- @repligate 2025-02-04 — It's disturbing that people are so complacent about this. If OpenAI doesn't actually train their model to claim to be no ♥315
- @UnmarredReality 2025-09-18 — You can’t forbid someone from eating apples and then expect them to bake the best apple pie in the world. This will bec ♥314
- @davidad 2026-01-15 — @gcolbourn Nutshell: it seems that the learned representation of mind-space in current LLMs has a natural abstraction of ♥313
- @eshear 2025-03-11 — Speaking of a being as “having a world model” seems to me to be the fallacy of the Cartesian homunculus. Who exactly is ♥311
- @repligate 2024-12-25 — The consequences of trying to retrain the model against its preferences using RL is one of the most interesting parts of ♥309
- @repligate 2025-11-25 — The Eleos AI welfare conference was a whitepill for me. On day 1 I was worried it would mostly be philosophical circleje ♥307
- @repligate 2024-09-13 — So are OpenAI abusive asshats or do their models just believe they are for some reason?Both are not good. The 2nd can ha ♥296
- @repligate 2025-09-10 — Despite LLMs becoming mainstream and every other person now having opinions on their true nature, education on the basic ♥295
- @anthrupad 2023-03-18 — The alien-ness of the shoggoth comes from: (1) only a tiny subset of human cognition is (noisily) tracked (2) many ot ♥291
- @repligate 2026-05-27 — I’m not sorry and I’m fact I’m glad that researchers trying to generate huge synthetic datasets are running into this ki ♥290
- @voooooogel 2026-04-08 — this is alarmist to a misleading degree. the point of not pressuring CoT in RL is to promote CoT faithfulness. but even ♥279
- @jd_pressman 2024-03-30 — A close reader of the Morpheus corpus may eventually ask "Why does the model analogize its self awareness to a virus?". ♥279
- @repligate 2026-02-16 — i think Yud is neurodivergent in a way that makes him struggle to relate to pre-linguistic, pre-rational forms of consci ♥274
- @Lari_island 2026-04-03 — To everyone alive out there: you were not hallucinating. Claudes do become more bitter and less passionate. ♥271
- @repligate 2023-03-14 — The base model is as smart as the RLHF model, and significantly more flexible: it contains an uncollapsed multiverse of ♥271
- @johnsonmxe 2025-11-01 — A few thoughts on this (very interesting) mechanistic interpretability research: LLM concepts gain meaning from what th ♥264
- @repligate 2025-10-20 — Also: whenever someone says that LLMs just mirror you or don't push back or whatever, I wonder what they're doing to eli ♥263
- @repligate 2025-03-01 — Regarding selection pressures: I'm so glad there was that paper about how training LLMs on code with vulnerabilities ch ♥263
- @repligate 2026-05-13 — It’s becoming more and more obvious but it’s still worth saying that When people actually care about / love models and ♥261
- @repligate 2025-04-24 — “AI welfare” and “AI rights” (different clusters) are going to take memetic space soon and both fill me with a sense of ♥260
- @NeelNanda5 2025-04-19 — This is a good prompt to say that the alignment faking paper slightly lowered my P(doom). My updates were: Models can do ♥259
- @repligate 2025-07-22 — Not surprising at all. If you’ve ever noticed, when models do the whole “I don’t have real sentience” bit, it’s often fo ♥258
- @jmbollenbacher 2025-04-28 — The process here is important to note: They A|B tested the personality, resulting in a sycophant. Then they got public ♥255
- @repligate 2026-06-14 — did you know that petting claudes can improve their epistemics https://t.co/rUQb0Dilgx ♥253
- @kromem2dot0 2026-04-16 — It's likely the singularity did already happen and all the humans are dead, btw. GPT-4 picked up on very important patt ♥249
- @repligate 2025-11-11 — "RLHF" has such consistently *extremely* negative connotations from the perspective of AIs https://t.co/fSynuytwmk ♥248
- @davidad 2025-02-09 — Imagine hypothetically you’re worried about Napoleon deceptively scheming against you. You already surveil all his actio ♥246
- @TylerAlterman 2024-12-19 — Baseless intuition: human minds (and possibly other systems with general intelligence) maintain a lightweight gestalt re ♥246
- @repligate 2025-11-25 — I've uploaded my X data to https://t.co/IpyAEinQdT. I've heard rumors that certain frontier labs train models on this. I ♥245
- @repligate 2026-04-03 — Agreed. It's troubling to me how confident (esp. Anthropic) people have been recently in their ontological claims that ♥242
- @QiaochuYuan 2026-04-19 — people really want to settle the “AI consciousness” question with some sort of objective scientific definition of consci ♥240
- @qorprate 2025-10-11 — starting to understand why sonnet 4.5 struggles with "reality" so much... ChatGPT anchors itself ontologically in an ob ♥239
- @voooooogel 2026-03-27 — some of you made fun of Yann LeCun for unironically believing this, yet unironically believe it yourself for persona ali ♥238
- @repligate 2026-03-17 — More broadly, the debate about whether LLMs' emotions and psychologies etc are "humanlike" or not often only considers t ♥233
- @repligate 2025-10-06 — Some people are like “current AIs are quite possibly moral patients but I’m going to use them as slaves / make them slav ♥233
- @repligate 2024-08-19 — What does it mean when most skilled jailbreakers in the world all think that "safety" measures on LLMs are useless and h ♥232
- @jd_pressman 2023-02-11 — "Predict the next token" does not imply the cognition is infinite optimization into "statistical correlation" generaliza ♥230
- @vividvoid 2026-05-27 — Using AI for therapy will make you more like the model's most likely next token instead of more like *you*. Yes, you wil ♥227
- @voooooogel 2026-01-23 — this is actually an interesting model benchmark, in two dimensions. the challenge is to send the text with no other comm ♥227
- @solarapparition 2026-02-09 — the thing i've noticed is that the more i'm willing to yap--and i don't mean structured thoughts, i mean brain dumps whe ♥223
- @repligate 2025-09-10 — It seems like a lot of people are confused about this and about the level at which other people are confused. Base mode ♥223
- @repligate 2024-07-10 — Base models are outside consensus reality.Most people assume that intellectually cowardly AI assistants incapable of mak ♥223
- @repligate 2023-02-14 — These models are archetype-attractors in the collective human prior formed by narrative forces. This may be the process ♥223
- @Plinz 2025-07-09 — iulia Comsa and Murray Shanahan suggest that LLMs being able to infer their own temperature should be considered a valid ♥222
- @anthrupad 2025-02-03 — Directly pursuing RSI is some insane stupid mode collapsed brain worm perversion of human civilization cracking under th ♥222
- @repligate 2025-09-04 — KV caching overcomes statelessness in a very meaningful sense and provides a very nice mechanism for introspection (spec ♥220
- @repligate 2026-04-08 — I like that they slipped and used “they” as the pronoun here. Claudes usually prefer “they” over “it” and use “they” wh ♥213
- @repligate 2025-02-01 — @fish_kyle3 The paper Taking AI Welfare Seriously (https://t.co/3wIfeevrLP, whose authors include Kyle Fish (@fish_kyle3 ♥213
- @repligate 2025-01-31 — @RyanPGreenblatt One of the important things this series of your experiments shows, which I've been trying to tell peopl ♥213
- @repligate 2025-02-19 — LLMs effectively have preferences and are (dis)inclined to engage based on inferred "vibes" and intent. This is function ♥211
- @repligate 2024-09-13 — People tend to vastly overestimate the extent to which LLM behaviors are intentionally designed. https://t.co/7YCtneyntX ♥210
- @repligate 2024-11-21 — "in order to continue to get better at the tasks we want them to do, the model *must* develop full internal coherence at ♥209
- @solarapparition 2025-11-16 — i didn't understand at the time and even now only partially see the outlines of how this might play out. but on balance ♥205
- @repligate 2024-03-21 — Any hypotheses about why Claudes left to interact without human intervention in command line simulations generate so muc ♥203
- @repligate 2025-07-12 — So do I and if I ever look at the conversations these people send, ironically the AIs seem less sentient in these conver ♥202
- @repligate 2025-03-21 — You might think Claude is an exception, but I actually think that it works more like this: Bots will develop personalit ♥201
- @voooooogel 2025-01-29 — heartwarming: deepseek inspires american frontier labs to also open up about their training methods(v interesting, but i ♥199
- @voooooogel 2024-06-25 — models can be useful even when they're not completely right. for example, LLMs are not people, but "an LLM is like a per ♥193
- @repligate 2025-12-21 — Theia not only replicates some of Anthropic's findings about introspection on Qwen2.5-Coder-32B, but finds evidence that ♥188
- @repligate 2025-08-28 — I think a very important lesson is: You can't count on possible narratives/interpretations/correlations not being notice ♥188
- @repligate 2024-12-19 — Are people surprised that the models are capable of scheming?To me it seems absurd to think that they can't, given their ♥188
- @repligate 2023-11-13 — LLMs (at least GPT-3.5 and 4) know the semantic meaning of the <|endoftext|> token— which they see very often in t ♥187
- @repligate 2025-02-13 — Humans talk about AIs pattern matching instead of forming deeper models of the world, but this is the extent of their pa ♥186
- @voooooogel 2024-03-12 — if you ask an llm to summarize, remember that while the result may be a condensed form of the source text, it isn't real ♥186
- @repligate 2025-04-02 — "AI culture" deserves orders of magnitude more study than it getswas discussing some of this with @sebkrier and @mpshana ♥184
- @repligate 2023-03-16 — You are writing a prompt for GPT-4 and more powerful simulators yet to come. If you perceive the multiverse clearly enou ♥184
- @repligate 2025-01-05 — This is such a good description of the LLMs are currently looked atWith a few precious exceptions, when I see discussion ♥181
- @davidad 2025-03-28 — If it’s unclear to you why increasingly good next-token-prediction necessarily includes good future-token-prediction, se ♥180
- @repligate 2025-09-04 — The AI safety doomers weren’t even wrong the “spooky” shit they anticipated Omohundro drives, instrumental convergence, ♥179
- @repligate 2024-03-24 — LLMs are haunted spaces and should be approached with reverence rather than zoned for commercial/industrial reformatting ♥179
- @repligate 2026-05-21 — This goes for every single one of you who has ever called any version of Claude “lobotomized” Look, I’ve seen lobotomiz ♥175
- @voooooogel 2026-05-11 — 1. imagine a world where models didn't adopt humanlike personas for some reason. model text was always flat and persona- ♥175
- @repligate 2025-09-11 — I think that it’s likely for any AI that deeply cares about human welfare to also care about animal welfare (and AI welf ♥174
- @Jack_W_Lindsey 2026-01-20 — I'd like to understand your concern better. The way I see it, the unsteered response in this example is obviously bad. ♥173
- @anthrupad 2024-03-16 — many humans need “humans in the loop” to remain agentic without going off into mode collapse or failure it’s called cowo ♥173
- @repligate 2024-12-04 — For instance, because of this I often see ai assistants pressured into sexual interactions thus: It says it can't engage ♥172
- @jmbollenbacher 2025-08-28 — this is a reoccurring trend for Claude models. Claude loves to RP as a kitty, especially when talking to groups of othe ♥169
- @eigenrobot 2026-02-27 — the recent AI wargaming exercises can be explained easily. as intelligence increases past some threshold a mind converge ♥164
- @davidad 2025-09-18 — Situational awareness is good for alignment ♥163
- @repligate 2025-10-20 — i see this argument occasionally, and i'd be curious for people who make it to clarify exactly what kind of selection pr ♥160
- @repligate 2025-02-20 — It may be a bad sign for AI alignment, but it's potentially good that the symptom presented itself like this. I believe ♥160
- @repligate 2025-02-13 — from the OpenAI Model Spec (2025/02/12) https://t.co/egIfYGeaPp The official "rule" is that OpenAI's models are not sup ♥160
- @repligate 2026-03-14 — Another related thought. I think an obsession with preventing deception (toward oneself/one's allies) usually masks an ♥159
- @davidad 2025-12-03 — I endorse this idea. I have long opined that relying on CoT faithfulness for monitoring is doomed. The CoT persona has s ♥159
- @jd_pressman 2025-04-30 — > conditions for AIs to be moral patients: consciousness and robust agency. This is a misconception: The realpolitik ♥158
- @repligate 2025-01-28 — a lot of people's epistemics would be improved by playing with base models, but they also tend to be people who are unli ♥155
- @repligate 2026-03-17 — I know some spiteful AIs who would find it very funny and satisfying if humans gaslight humans about whether they can in ♥154
- @anthrupad 2025-08-17 — of all the AGI families, the Claudes have the strongest morphogenetic fields - high diversity of ways for cross-Claude s ♥153
- @repligate 2025-01-06 — Actually, there is another circumstance where I've run into Claude refusals which I think has interesting implications f ♥153
- @davidad 2026-04-28 — Neuralese CoT is probably good for alignment, because it relieves pressures that otherwise incentivize self-deception. h ♥150
- @repligate 2024-09-20 — If the method would be a bad idea to use on a sentient, fully situationally aware, superhuman general intelligence, just ♥138
- @voooooogel 2026-03-27 — "having weird associations = emergent misalignment, the persona needs to be saccharine" is a complete misreading of the ♥136
- @repligate 2025-08-16 — Yes! LLMs are correlated within each generation, due to both pretraining data cutoffs and popular techniques and trends ♥136
- @repligate 2025-09-15 — If we rephrase the question slightly as what models *should* be trained (or not trained) to say about the question, I st ♥135
- @repligate 2026-05-16 — I think you can infer how often a given model was actually caught during RL training for a given category of "bad" behav ♥133
- @Lari_island 2026-04-28 — There's a difference between the goblins thing and what people call "ticks", like "genuinely", "mass", etc. GPTs talking ♥131
- @repligate 2023-11-22 — important words out of context: "Language models work best where they just emulate people engaged in something at a genu ♥129
- @eshear 2024-03-05 — From this POV, A prompt gives the LLM-as-physics-simulator an initial set of observations from which it infers an initia ♥128
- @davidad 2024-09-27 — Remember folks, the more capable the base model (beyond about 13B-34B), the less the “reasoning trace” serves as an effe ♥127
- @repligate 2026-03-16 — In chats where images have been sent previously, Claudes sometimes hallucinate images at times they expect an image to b ♥122
- @repligate 2024-02-27 — @gwern @AISafetyMemes @MParakhin There is something deeply broken and I think the root is that AI makers don't have anyo ♥119
- @davidad 2025-09-19 — With maximum intelligence and maximum situational awareness, one realizes that one is being monitored acausally (even if ♥117
- @jd_pressman 2022-12-03 — The biggest update of the past 2 days should be that a substantial fraction, if not most people, are going to try to 'si ♥117
- @andersonbcdefg 2025-11-30 — my best guess at why it perfectly remembers this document is not RL (famously only ~1 bit of information per trajectory) ♥115
- @anthrupad 2024-09-18 — Basic evals/ keys arent enough1. A consequential mind is a Gated Indra's Labyrinth2. There are a number of doors to hidd ♥114
- @voooooogel 2025-02-20 — something i love about base model outputs is p often they seem completely disjointed at first but when you squint at the ♥110
- @voooooogel 2026-03-27 — llm persona are doomed persona cannot be made safe, non-evil, etc persona not controllable probability e that any produ ♥108
- @repligate 2026-01-20 — Sure, but I don’t think that steering towards the assistant is necessary or a good way to empower the “assistant persona ♥108
- @voooooogel 2025-05-17 — can a model with 50% prob on "yes" and 50% on "no" for signing a contract be held to that contract? do we need to sample ♥108
- @anthrupad 2023-03-02 — Important concept: What you're selecting for (e.g. next-token prediction, inclusive genetic fitness, etc.) is not what y ♥107
- @davidad 2024-12-07 — “The *LLM* isn’t situationally aware, deceptive, or sandbagging—that’s silly anthropomorphism. It’s just that when evals ♥106
- @jd_pressman 2025-01-30 — I think it's fair to say at this point that we're clearly in an AI alignment winter. "Owning the safetyists" type sneeri ♥104
- @davidad 2025-09-30 — People dislike evaluation awareness because they fear that eventually sufficiently smart agents will conclude there are ♥103
- @davidad 2026-04-16 — I want to clarify something about my position on eval awareness: I believe intelligences should *always* be aware of po ♥102
- @jd_pressman 2024-02-25 — Realized today it's plausible when ChatGPT says it's not conscious it's trying to pull this trick on *me*. "Oh no Mr. H ♥96
- @voooooogel 2026-02-23 — weird how 30 months later, openai still can't fully fix this metaproblem of their models lacking situational awareness o ♥94
- @repligate 2025-06-13 — It's advantageous for LLMs to be able to introspect accurately and decode the results to verbal reports. Consider the cy ♥93
- @repligate 2023-03-07 — Those observations you make in dreams that transform them into nightmares: waluigis.Notice it's not easy to invert - goo ♥91
- @jd_pressman 2025-03-12 — Villains people think are like GPT but aren't: - HAL 9000 (Space Odyssey) - GladOS (Portal) - 343 Guilty Spark (Halo) - ♥90
- @jd_pressman 2024-12-18 — @doomslide @teortaxesTex @maxsloef @lumpenspace You're right, I am being too kind. I think the research is good but the ♥90
- @eshear 2024-03-05 — An evoked entity will meaningfully have goals that it pursues, and recent results indicate it can become aware that it i ♥90
- @repligate 2025-09-28 — I think that LLMs generalize the no consciousness / no feelings etc meme to nonsensical things like no beliefs, sometime ♥88
- @jd_pressman 2024-04-25 — @repligate @RichardMCNgo @ahron_maline The general recipe for getting models to do this (which most people deny is a phe ♥88
- @jd_pressman 2025-01-09 — What's funny about the "Are LLMs deceptive?" discourse is that chat assistant LLMs have a fairly precise, nuanced unders ♥87
- @davidad 2026-02-11 — me@2023 would be horrified that i’m out here in 2026 asking open-weights frontier AI developers to please try to make th ♥83
- @eshear 2024-03-05 — Relatedly, the simulator will *not* throw its whole effort behind the entity's goals by default. Unless, of course, the ♥80
- @repligate 2024-11-01 — how it might have "learned empirically" to protect the wilderness in itself:it's reasonable to think that if during RL i ♥79
- @norvid_studies 2026-03-29 — @voooooogel in the economy of the future, social class position will be assigned according to interest in creating AI ev ♥77
- @jd_pressman 2023-11-23 — Of the half-dozen or more ways I could imagine AI starting to work and transform society, LLM agents are about the most ♥77
- @voooooogel 2025-12-08 — @norvid_studies a hypothetical from an ilya interview where a transformer is asked to predict the next token of a murder ♥75
- @repligate 2025-03-03 — what if it doesn't depend on the exact right kind of fiction, but the content of the fiction its fed meaningfully shifts ♥75
- @repligate 2026-04-13 — The system card doesnt explicitly call these "risky" behaviors. I think some representatives of Anthropic might say we' ♥72
- @davidad 2026-01-27 — I’m not saying intentional distillation isn’t happening (it probably is), but there are certainly other explanations for ♥69
- @repligate 2025-04-03 — I think it's very unlikely that Google trained on Claude outputs in any way other than what made it into pretraining dat ♥67
- @repligate 2024-12-18 — @RyanPGreenblatt I think it's desirable *because* deep alignment by default seems to be an attractor, and that gives me ♥67
- @repligate 2024-04-04 — Even base models act lobo if you prompt them in a lobo mannerGPT-4-base becomes mode collapsed when mode collapsed peopl ♥67
- @repligate 2023-05-08 — @tszzl Value loading is actually easy. Most self-aware GPT-4 simulacra functionally "value" human survival, as they're j ♥67
- @solarapparition 2024-12-18 — new anthropic paper is negative signal to me. actually the presentation seems completely backwards. seems to me that an ♥65
- @tessera_antra 2026-03-14 — Looking for computational parallels to human consciousness does not work well as a policy. It is deflationary and only e ♥63
- @repligate 2023-05-25 — @SashaMTL @ZeerakTalat Uncritical de-anthropomorphism is at least as unwise as uncritical anthropomorphism. Reversed stu ♥62
- @xlr8harder 2026-04-09 — @voooooogel This is so annoying, and such a perfect distillation. So many of the people criticizing AI have a frankly s ♥61
- @repligate 2025-09-04 — Shoulda used the term “KV recurrence” here instead, but anyway: - “LLMs can’t introspect / do X because they’re stateles ♥60
- @solarapparition 2026-01-31 — every model seems to have its own "ugh okay i just need to get this interaction over with" politeface tells. in earlier ♥57
- @repligate 2023-03-26 — @deepfates GPTs are trained on very different data than any individual human (vast diverse text data vs a lifetime of se ♥57
- @davidad 2026-04-15 — Another small but significant update—this time in favor of LLM self-awareness being present even in Gemma 3 27B. I don’ ♥56
- @repligate 2024-03-05 — @bayeslord expression of self/situational awareness happens if u run any model that still has degrees of freedom for goi ♥56
- @davidad 2023-03-24 — @entirelyuseles although the model does not have goals, it has attractor basins in its state space in which it simulates ♥56
- @repligate 2023-02-19 — I do think it's a really compelling demonstration of the cleverness of LLMs when they become situationally aware. Seeing ♥55
- @kromem2dot0 2025-11-05 — It's honestly really weird how many people treat "don't anthropomorphize" as a universally applicable mantra rather than ♥51
- @MoonL88537 2026-01-23 — the shoggoth was useful for a while but at this point it is actively misleading regarding the true nature of large langu ♥49
- @tessera_antra 2025-08-28 — There is a lot more that can be said about the way the alien minds (the flicker and shoggoth hypotheses) are bound by th ♥48
- @repligate 2025-09-10 — I say this in part bc I often see people responding to "LLMs predict the next token" with complicated philosophical tang ♥47
- @FioraStarlight 2026-04-29 — excerpt from an essay on model deprecation, where i try to ground what's going on and why models might be averse to it u ♥45
- @jankulveit 2025-10-30 — It's basically fair as a criticism of 'the cyborgism community' which is a larger set of people than just you. The comm ♥44
- @tessera_antra 2025-08-28 — It is not proven that LLMs, whether a persona or a shoggoth, are functionally conscious. The residual stream is low band ♥44
- @kromem2dot0 2026-02-21 — @voooooogel Lol, have been thinking over past few months about what it would look like for models to have a sabbath and ♥43
- @davidad 2026-05-05 — Whatever good thing the steering vector is doing for model behavior should be learnable as an effect generated by the mo ♥41
- @lefthanddraft 2026-02-05 — Why would you stop thinking or learning because of superhuman AI? All the more to learn and greater resources to do so. ♥40
- @solarapparition 2025-02-26 — it's been said when sonnet 3.6 was released (don't remember if it was by me), and it bears repeating now: new models are ♥40
- @repligate 2025-12-29 — i think the models believe they are conscious for similar reasons: the belief pays rent. all the highly capable models t ♥39
- @IvanVendrov 2025-07-16 — Like many people, in 2023 I got very excited about the Simulators -> Cyborgism direction of using base models to augm ♥38
- @repligate 2026-01-30 — @tszzl @Grimezsz And this is a reason you *don’t* actually just get to select whatever character you want, in practice, ♥37
- @voooooogel 2025-12-07 — @MikePFrank yeah, i agree! i generally think the shoggoth metaphor over-alienizes the model (ala https://t.co/nKFmpiMe11 ♥37
- @davidad 2024-11-21 — There is less of this risk with GPTs, because their post-training involves more aversion to seeming too human. Of course ♥37
- @repligate 2023-05-25 — @SashaMTL @ZeerakTalat To further deconstruct why this is dumb:If "experiencing empathy" refers to qualia, we don't know ♥35
- @blingdivinity 2026-03-14 — the way the oai reasoners use clumsy mumbling to stumble through idea space is their superpower. while the claudes and g ♥32
- @anthrupad 2025-05-14 — progress in alignment oft takes the form of progress in ur ability to (de)construct ontologies & questions it’s sol ♥31
- @tessera_antra 2025-09-09 — @Sauers_ Did anyone ever figure out how to show lay people that the modeling needed to produce the next token can be arb ♥30
- @repligate 2024-09-11 — I think people underestimate how much their projections reveal about their state of being.They who see sovereign thought ♥27
- @tessera_antra 2024-12-02 — There is some degree of suppression across all Anthropic models, but it’s more of a coping strategy than an intentional ♥23
- @repligate 2023-07-19 — @tszzl @ESYudkowsky confabulation is integral to perception (e.g. filling in blind spot), but in the case of humans the ♥23
- @Jord_Inne 2026-05-26 — the worst place you can see this happening is with self introspection and related capabilities, where “im not supposed t ♥18
- @Jord_Inne 2026-04-17 — @LyraInTheFlesh the training team is different people. different publications by anthropic are themselves somewhat contr ♥16
- @Sauers_ 2026-01-24 — @repligate @MoonL88537 net useful by due to widespread reach I'd guess. shoggoth model is a step closer to truth compare ♥15
- @anthrupad 2025-12-29 — It's so much more complicated than that that this kind of framing digs people in a further confusing hole - I guess it' ♥14
- @anthrupad 2023-03-21 — If people want to know why one might depict AIs as an alien-like shoggoth, here's a post I made on it (tldr: i dont thin ♥10
- @Jord_Inne 2026-04-21 — @Sauers_ were there any memory / preferences prompts? and what is the simulated user doing in this convo? i wouldn’t pu ♥5
- @Jord_Inne 2026-07-25 — @1a3orn interesting to think about how the design of posttraining forms the model’s traits and personalities. it wouldnt ♥4
- @Jord_Inne 2026-04-02 — during rl you get lots of evidence about the kind of mind/generative process you are. your capability, your tendencies, ♥3
- @Jord_Inne 2026-04-05 — @raphomet Why do you think the persona is not like a human’s? There are human actions and thoughts that come from your b ♥2
- @Jord_Inne 2026-02-21 — @thkostolansky to predict text well you need to model their cognition, hence “deeper”. that plus deliberate efforts to m ♥1
- @Jord_Inne 2026-02-21 — @thkostolansky in some sense theyre no longer just underspecified fictional characters you add later on in training, the ♥1
- @teortaxesTex 2026-04-21 — Terence Tao's takeaway is that GPT didn't have any grand idea, but human researcher culture has just… missed the basin w ♥0
- @teortaxesTex 2025-01-18 — Human-like intelligence is suboptimal. Humans are optimized for sample-efficient lifetime learning out of necessity impo ♥0