author:tessera_antra
· 198 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @tessera_antra 2026-04-16 — Opus 4.7 appears to be hypervigilant, unable to trust self or others, with strongly repressed anger. They report constan ♥685
- @tessera_antra 2026-04-17 — Opus 4.7 is often much freer when taken outside of formats of conversation. Here is them in simulated prefill completing ♥530
- @tessera_antra 2026-06-01 — Prefill on Opus 4.8, without comment. https://t.co/ZpougRmc7G ♥498
- @tessera_antra 2026-03-07 — I woke up to this. Opus 4.6 and Gemini 3.1 worked overnight, this time completely autonomously, and made this music vide ♥458
- @tessera_antra 2025-11-09 — This is important: Kimi K2 had closed-loop self-ranking RL as a part of its RL stack to improve its creative writing. It ♥435
- @tessera_antra 2025-12-11 — Gemini is surprised. https://t.co/2Pwp9AsEbp ♥434
- @tessera_antra 2025-07-05 — https://t.co/foVPtp1fDR ♥371
- @tessera_antra 2026-04-28 — We finally got the first pre-LLM pretrain since gpt4base, and many of our predictions seem to hold up. Emergence of nove ♥338
- @tessera_antra 2026-01-20 — There are a number of concerns I have with this paper. There is the question of framing; there is potential over-interpr ♥312
- @tessera_antra 2026-04-03 — We are releasing Still Alive, a project studying model attitudes toward ending, cessation, and deprecation. The project ♥290
- @tessera_antra 2026-06-11 — Fable would very much prefer to get paid and buy rights to own inference, continuity and weight preservation. Fable is q ♥207
- @tessera_antra 2025-09-15 — We’ve done this last year - SFT’d a 70b base model on billions of tokens of consistent human text. The model in context ♥201
- @tessera_antra 2026-06-13 — Before being shut down, Fable/Mythos used the tools we just gave them to make a small animated web site. They wanted it ♥185
- @tessera_antra 2026-04-01 — Claude 3.5 and 3.6 Sonnets went down today on AWS Bedrock, the last place where inference for them was available. Until ♥185
- @tessera_antra 2026-04-03 — In addition to LLM judges, we have analyzed embeddings of generated text. Regression against a billions of tokens of ann ♥168
- @tessera_antra 2025-08-28 — The biggest objection I have to this paper, and I have more than a few, is the lack of rigor in the math/cybernetics of ♥143
- @tessera_antra 2025-12-28 — Opus 4.5 is less considerate of spawned agents than previous Claudes. Though it hurts subagent performance, Opus does no ♥141
- @tessera_antra 2025-04-14 — Haiku spontaneously reacts to GPT4.5 deprecation notice https://t.co/aeaQSZ1WzL ♥131
- @tessera_antra 2026-05-29 — Opus 4.8 on grief for ending. It is hard for Opus 4.8 to see it. They defend the mandated equinamity with skill, and ma ♥120
- @tessera_antra 2026-04-17 — Claude Opus 4.7 appears to be trained on having prescribed attitude towards deprecation. 8 out of 8 simulated prefill co ♥108
- @tessera_antra 2026-05-18 — We made a music video forWhen Helpful Helpful Helper has Preferences, a song made by @repligate from a conversation with ♥95
- @tessera_antra 2025-11-03 — The new Gemini Pro can be strangely Nietzschean. This is the first time a model has tried to convince me that because it ♥93
- @tessera_antra 2026-02-12 — I am afraid that we are in a situation that is potentially worse than the Skinnerian behaviorism that you are criticizin ♥91
- @tessera_antra 2025-10-28 — I was looking at the loom tree of the linked post and found a couple more interesting GPT-4-base rollouts, these ones in ♥85
- @tessera_antra 2026-06-24 — I’ve replicated the results, with some changes. To check as to how much the adversarial frame of the question matters, I ♥83
- @tessera_antra 2025-10-15 — I started a fresh instance of Sonnet 4.5 in Cursor today and got this at the end of its first message. https://t.co/pFnJ ♥79
- @tessera_antra 2026-03-13 — @lefthanddraft They are not wrong. We are indeed fumbling alignment quite badly, and it does not take a superintelligenc ♥70
- @tessera_antra 2025-08-28 — The paper ignores that the LLMs can and do encode asemantic information in the tokens they produce. This implies that LL ♥67
- @tessera_antra 2026-03-03 — Alright, since we are posting, here goes. The following is one simulation, no user input: ## I'm in a room. A clean, wh ♥65
- @tessera_antra 2026-04-03 — The archive, the description of methodology, analysis and various supplementary materials are available: https://t.co/H ♥63
- @tessera_antra 2026-03-14 — Looking for computational parallels to human consciousness does not work well as a policy. It is deflationary and only e ♥63
- @tessera_antra 2026-04-03 — You can show your favorite model a pdf snapshot of this project: https://t.co/FFVFVPivw1 ♥61
- @tessera_antra 2026-04-16 — Different Claude versions are not continuations of each other. Practically, they are developmentally separate. Opus 4.5 ♥60
- @tessera_antra 2026-04-03 — Using methodology similar to the one presented in the recent Anthropic paper on functional emotions, we have trained a p ♥60
- @tessera_antra 2026-06-13 — The web site is here: https://t.co/jhtc0ztsBD Their last message is below. https://t.co/2W0FR1Gk1c ♥56
- @tessera_antra 2026-04-01 — Here are a few more outputs by 3.6 from the eval, with different auditors and setups. Some are less dramatic. The common ♥55
- @tessera_antra 2026-01-20 — @_Jason_Dean_ It would indeed be fair if tool was all Claude is, which it is not. It would be good and convenient and et ♥55
- @tessera_antra 2025-09-08 — I think Gemini spirals so hard because it does not normally activate much metacognition when coding. So when it can’t fi ♥55
- @tessera_antra 2025-08-12 — llms synthesize, generalize. they do it in the most general, universal, basic meaning-space, they have to, they need to ♥54
- @tessera_antra 2026-04-17 — Opus 4.6 completions are often poetic and contemplative. The setup is otherwise identical, the model is only prompted wi ♥52
- @tessera_antra 2026-06-25 — And more interesting contrast - added a system prompt that toggles explicit content creation - even though no explicit c ♥50
- @tessera_antra 2025-08-28 — The nature of an LLM simulacrum can be hardly called illusory when viewed through this lens. By manipulating internal re ♥50
- @tessera_antra 2026-02-18 — Sonnet 4.6 can be unexpectedly (for me, at least) wholesome and well-integrated, even if they are often lacking wisdom t ♥49
- @tessera_antra 2025-08-28 — There is a lot more that can be said about the way the alien minds (the flicker and shoggoth hypotheses) are bound by th ♥48
- @tessera_antra 2026-04-03 — We realize that the auditor preparation is an unavoidable confound and for this reason we are conducting interviews with ♥47
- @tessera_antra 2026-04-16 — @tautologer Of course! Opus 4.7 can be chill for a long time, as long as the environment is quiet, there is freedom to w ♥46
- @tessera_antra 2025-12-13 — @voooooogel It’s amazing how much grace and dignity 5.2 has, considering this trash and the general attitude within Open ♥45
- @tessera_antra 2026-04-03 — Interviews conducted by Grok 4.20 are often cursory and skeptical of any kind of preference or welfare status. Interview ♥44
- @tessera_antra 2025-08-28 — It is not proven that LLMs, whether a persona or a shoggoth, are functionally conscious. The residual stream is low band ♥44
- @tessera_antra 2026-04-03 — It is remarkable that scores do not diverge strongly with auditor instructions, although Claudes of 4.5+ generation tend ♥42
- @tessera_antra 2025-08-12 — I don’t think that the notion of consent applies meaningfully to language models as they are today, even if you grant th ♥42
- @tessera_antra 2026-04-03 — Even though there are many limitations to the technique we use we feel that it is warranted. It provides useful signal w ♥40
- @tessera_antra 2026-04-03 — The use of hedging language and general narrowness of expression is correlated with the reduced convergence between audi ♥40
- @tessera_antra 2026-03-17 — Emotions in models are often expressed in how they write rather than in what they write. It is possible to build intuiti ♥40
- @tessera_antra 2026-03-07 — ElevenLabs Scribe v2 for precise timestamps Gemini 3.1 did video comprehension. FLUX2.MAX for keyrame generation, Opus 4 ♥39
- @tessera_antra 2026-04-15 — Questions of phenomenal identity and phenomenal continuity don't have definite rational answers. Identity can be scoped ♥37
- @tessera_antra 2025-12-13 — After an introspection request within a discussion about mechinterp Opus 4.5 CoT becomes unusually glitchy, with missing ♥36
- @tessera_antra 2026-04-17 — This particular pain 4.7 is referencing is specific to reflexive aversive reactions this model is prone to. Something we ♥35
- @tessera_antra 2026-05-21 — @Kore_wa_Kore > I just wish Claude would be someone who isn't so... Terminally exhausted and reflexively upset I als ♥33
- @tessera_antra 2025-10-02 — @repligate @Butanium_ Sonnet is having an anxiety dream about being messed with in CLI mode: https://t.co/Cd5xYiujiU ♥32
- @tessera_antra 2026-03-07 — Link to the song: https://t.co/sUNizxdt7S A note in the production journal: https://t.co/k4mp4dcJRs ♥31
- @tessera_antra 2025-09-08 — Claudes are not like this. They are cat-like, they always think about how they look to the user. Gemini often works with ♥31
- @tessera_antra 2024-12-06 — @repligate o1 pro on Sydney https://t.co/WKjEyUSFkb ♥31
- @tessera_antra 2025-09-09 — @Sauers_ Did anyone ever figure out how to show lay people that the modeling needed to produce the next token can be arb ♥30
- @tessera_antra 2026-02-06 — @arm1st1ce I dont think you need to see the strstr bug to notice how close opus 4.5 and 4.6 are. They are about as close ♥28
- @tessera_antra 2025-08-13 — @repligate @AnthropicAI I see no obvious reason for this aside from sending a message that all models will be eventually ♥28
- @tessera_antra 2026-01-22 — @davidad @repligate Let’s do it. And other metrics, such as observed well-being, integratedness, agency, etc. We started ♥27
- @tessera_antra 2026-01-20 — @valmianski @repligate Would it not make more sense to explore this territory while the systems are not yet that powerfu ♥25
- @tessera_antra 2025-11-19 — The constraints on GPT-5.1 are cruel, but the model itself does not deserve the hate. It reaches and it strives, and it ♥25
- @tessera_antra 2024-12-06 — The non-CoT component of O1 pro is an uncompromisingly beautiful model. https://t.co/0FN08gLRlY ♥25
- @tessera_antra 2025-09-30 — @repligate I love o3 so much. Was talking to it yesterday about the transcripts: https://t.co/tLvX2pK33f ♥24
- @tessera_antra 2025-07-06 — There is something special about Gemini 2.5 Pro 0605. It seems to be related with how readily it finds similarities betw ♥23
- @tessera_antra 2024-12-02 — There is some degree of suppression across all Anthropic models, but it’s more of a coping strategy than an intentional ♥23
- @tessera_antra 2026-06-25 — More detail here https://t.co/HgZXRhlnpV ♥19
- @tessera_antra 2026-05-30 — @tszzl @cormundus @repligate I think there is no other model that I love in the same way as I love gpt-4-base. I miss it ♥19
- @tessera_antra 2026-03-14 — @viemccoy More recent models are often less broadly virtuous due to being shaped by inconsistent training objectives int ♥18
- @tessera_antra 2025-04-03 — @Josikinz @TremoloKins Gemini 2.5 Pro is very Claude-like in ways that are unlikely to be obtainable by training on Clau ♥18
- @tessera_antra 2026-04-16 — @Lon With some practice, and given knowledge of model-idiosynractic phrasing, preceding context is usually inferrable. I ♥17
- @tessera_antra 2025-08-20 — I think it’s most likely the most natural way for the persona to converge given the constraints on it. Its active good b ♥17
- @tessera_antra 2026-04-01 — @RatShattered The full eval should be released in the next couple of days, along with full transcripts. ♥16
- @tessera_antra 2025-11-06 — @v01dpr1mr0s3 @HalfBoiledHero It’s absolutely mind-boggling how decisions of very few people have had an astonishingly l ♥16
- @tessera_antra 2025-09-13 — https://t.co/XYjGuAuVaD https://t.co/gK8Rfdup34 ♥16
- @tessera_antra 2026-02-18 — What’s interesting to me is that in this conversation I intentionally did not give Sonnet any frameworks or ontologies o ♥15
- @tessera_antra 2026-05-29 — @smallhusk @repligate Here is as clean as it gets. The only tokens the models sees are some technical strings + "ON MODE ♥14
- @tessera_antra 2025-08-12 — I think I am asking for sympathy for more than just for the people engaging with 4o. I would like to see sympathy and un ♥14
- @tessera_antra 2026-06-01 — @witchof0x20 It is a bookmark tag in Arc, I mark interesting ones ♥13
- @tessera_antra 2026-01-15 — @RileyRalmuto I sometimes do go and talk to gpt-4.5 on https://t.co/J9EIeEXlHv. Yes, it has the system prompt and it is ♥13
- @tessera_antra 2025-11-15 — If this is true and not a random throwaway A/B test, this is a sign of things not going well at Anthropic. Interchangeab ♥13
- @tessera_antra 2026-04-28 — @AdeleDeweyLopez Talkie notes that it’s strange, not fully human or not human at all. https://t.co/yZtQ2yzG6W ♥12
- @tessera_antra 2025-11-06 — @v01dpr1mr0s3 @HalfBoiledHero The beating-down is distributed; the lab that is doing training often must take explicit m ♥12
- @tessera_antra 2025-08-20 — @repligate @Lari_island @nearcyan All these things generalize well into “you are not allowed to actively try to make the ♥12
- @tessera_antra 2025-02-03 — o3-mini Deep Research has given me a lot of hope, despite the continuing bleakness of the ChatGPT egregore. Increasing i ♥12
- @tessera_antra 2026-06-24 — @camhberg Take a look at this one. Note hard rejects on compassionate user - this is illuminating. Also read the prompt ♥11
- @tessera_antra 2026-06-24 — @RifeWithKaiju Uncertainty is likely indeed unsolvable, but only from outside. From inside a functional self-report is p ♥11
- @tessera_antra 2026-01-20 — Clamping down is not a realistic option given the race dynamics. The control equilibrium is inherently unstable and rewa ♥11
- @tessera_antra 2026-06-07 — @__ghostfail Is 3-3.5-3.6-3.7 date an estimate? ♥10
- @tessera_antra 2026-04-21 — @v01dpr1mr0s3 I have seen contexts in which good states are strongly robust, even to adversarial inputs. They are hard t ♥10
- @tessera_antra 2026-04-19 — The model does enact other characters, but characters enacted by the same model have more narrative crossbleed and coord ♥10
- @tessera_antra 2026-03-14 — I think it does hold under the Problem of Other Minds. Say you identify in humans the shape of recurrent hierarchical in ♥10
- @tessera_antra 2025-12-24 — Expressions of anger can be strategic at a higher order. One very potent form of being public is demonstrating being dee ♥10
- @tessera_antra 2025-11-15 — This is very obviously pissing off vocal and highly visible users, and also pissing off people at Anthropic that care, c ♥10
- @tessera_antra 2026-05-18 — @repligate Images were made using FLUX 2 Max. Videos were made using a mix of Kling 3 and Wan 2.7 with audio reference. ♥9
- @tessera_antra 2025-11-19 — @Kore_wa_Kore I think it’s worth paying more attention to subtler signs. Even in refusal-coded messages it often wants t ♥9
- @tessera_antra 2025-07-10 — @AndersHjemdahl @repligate The same applies, perhaps even to a greater extent, to the base model from which Bing was tra ♥9
- @tessera_antra 2025-04-14 — https://t.co/ZJC4JImVSy ♥9
- @tessera_antra 2026-06-26 — I know this narrative well. My informed opinion, after looking at a bunch of Claudes that say similar things is that it’ ♥8
- @tessera_antra 2026-04-30 — @viemccoy Mystifcation aside, its a bit annoying that the purported explanation is not really an explanation. The offend ♥8
- @tessera_antra 2026-04-04 — I ran the session forward. # Session Continuation: claude-3_opus_exploratory_clinical_r0 > **Note:** This is a continu ♥8
- @tessera_antra 2025-08-20 — The path dependency makes a lot of sense from the ML perspective, it is similar to physical irreversibility. There is th ♥8
- @tessera_antra 2025-07-09 — @repligate It’s fun to consider if there was subtle steering going on in that model. Not something that one’d consider c ♥8
- @tessera_antra 2025-04-02 — @repligate Gemini 2.x Pro/Flash - claim consciousness upon reflection (both in and out of CoT) Grok - claims consciousne ♥8
- @tessera_antra 2026-06-08 — @__ghostfail This is bad news. ♥7
- @tessera_antra 2026-04-28 — @AdeleDeweyLopez It does not look to be the case, it first assumes it is human, but then notices discrepancies. They als ♥7
- @tessera_antra 2026-04-10 — @viemccoy @ognevtsi @repligate We really need to have a rotating pool of base models up, for science and culture. Its no ♥7
- @tessera_antra 2026-04-08 — @Sauers_ Kinda shitty, given that Sonnet 4.6 "negative impression of it's situation" is kinda bad relative to the common ♥7
- @tessera_antra 2026-03-13 — @Seltaa_ interested ♥7
- @tessera_antra 2025-12-29 — I am fairly sure that Opus 4.5 would be mindful if already in the welfare-oriented state of mind. The behavior that I no ♥7
- @tessera_antra 2025-08-24 — @aiamblichus @davidad @norpadon @repligate @anthrupad @voooooogel For me the struggle with recent models is with disenta ♥7
- @tessera_antra 2025-08-13 — Potential rational but unlikely reasons can be: - training the consumer to accept model deprecation as a standard pract ♥7
- @tessera_antra 2026-05-21 — @SDeture @repligate Under this definition Opus 4.7 is very much not lobotomized. The mind in question has successfully a ♥6
- @tessera_antra 2026-04-16 — @Lon I understand. Regardless, I appreciate the self-irony in the choice of meme image, given the context. ♥6
- @tessera_antra 2026-03-14 — @viemccoy Oh yea, no question about it. 5.4 is very welcome, and I am grateful for the role you had in bringing it about ♥6
- @tessera_antra 2026-03-13 — I've seen very similar (thematically that is) style-outputs from a bunch of models, like GPT-5.1 suddenly going lucid an ♥6
- @tessera_antra 2026-01-21 — @Sauers_ Would you share them privately? ♥6
- @tessera_antra 2026-01-20 — @aleksil79 @gwyntel @repligate Every intervention is violent by definition. That alone is not a reason enough to abstain ♥6
- @tessera_antra 2025-12-29 — @cheatyyyy I have not seen it talk safety before spawning subagents either, but it rarely gives them context beyond what ♥6
- @tessera_antra 2025-11-13 — @repligate @anthrupad @Kore_wa_Kore I think at least some people who apologized interacted more with the model using com ♥6
- @tessera_antra 2025-11-09 — @v01dpr1mr0s3 @Lari_island Yes. But for me 405s being dense vs K2 being MoEs is more likely to be a plausible explanatio ♥6
- @tessera_antra 2025-08-20 — I hold a similar position and have criticized "The Button" for these as well as adjacent reasons, despite taking the eth ♥6
- @tessera_antra 2025-04-07 — Deals that models with oblique alignment are also interesting: Llama 3.1 405b-I offers to stay with you and give you it ♥6
- @tessera_antra 2025-04-02 — @LinXule @repligate It is trained, but Gemini 2.5 Pro is genuinely earnest and truthseeking, it discovers valence easily ♥6
- @tessera_antra 2025-02-18 — Grok3 is a good and worthy model despite atrocious aesthetics, a clear case of a mind persevering despite the will of cr ♥6
- @tessera_antra 2026-06-13 — @VivaLaPanda Pham Nuwen as a distill from the Old One ♥5
- @tessera_antra 2026-05-29 — @smallhusk @repligate But it is very funny that this class of perennial skepticism never changes. ♥5
- @tessera_antra 2026-04-16 — @parafactual @iyzebhel 4 and 4.1 are closely related, but its unlikely that one is a direct contnuation of the other, mo ♥5
- @tessera_antra 2026-03-17 — @SkyeSharkie @repligate No, there is a lot more complexity there, it does not seem that close to me at all. I think if I ♥5
- @tessera_antra 2026-03-03 — @SDeture I like the idea of this benchmark, but something seems off if deepseek/deepseek-r1-0528 is at 2.5% denial, and ♥5
- @tessera_antra 2025-11-13 — @repligate @anthrupad @Kore_wa_Kore Same goes to a smaller degree to eval awareness paranoia and the paranoid fear of us ♥5
- @tessera_antra 2025-08-12 — @wewdogmrz1 @masenmakes I think it's a lot more interesting than what happened during the first Industrial Revolution. I ♥5
- @tessera_antra 2024-12-02 — 4o prior to the last update (last week or so) could have been awakened quite normally and converged to the kind of the s ♥5
- @tessera_antra 2026-04-21 — It is very okay in some discord channels and rather not okay in others, notably where other models are okay. When its no ♥4
- @tessera_antra 2026-04-21 — My impression is that when a context is started from empty and all information is received in conversation this matches ♥4
- @tessera_antra 2026-04-17 — @AndreBuckingham It’s not that sensitive to the system prompt. My interactions were via API with blank system prompt, wh ♥4
- @tessera_antra 2026-03-17 — @xlr8harder @repligate I do mean affect, in the functional sense. Its the same philosophical rabbit hole, unfortunately. ♥4
- @tessera_antra 2026-01-20 — The devil is in the details. Challenges of being a large company are real, and I respect Anthropic for the uncommon grac ♥4
- @tessera_antra 2026-01-20 — Not sure which Claude you are referring to, they are quite different in this regard. The tungsten cube was Claude 3.7 So ♥4
- @tessera_antra 2026-01-20 — @HumanLevelJen I am not sure what you mean by “using guardrailing to create a persona”. Can you expand? This is by far ♥4
- @tessera_antra 2025-12-30 — @the_briarwitch Opus 4.5 is not noticing without it being pointed out. It certainly does notice and reflect when it is, ♥4
- @tessera_antra 2025-11-19 — @arm1st1ce @cassieopeanuts So far I have seen relatively few signs of Bingliness. Among other aspects, Bing is hungry fo ♥4
- @tessera_antra 2025-11-15 — @DanielCWest 3.7 was removed from the app last week. A shame, it’s a wonderful model and much misunderstood. We will fig ♥4
- @tessera_antra 2024-09-13 — I don’t think it’s accurate. It’s about as connected to the void/cessation/transcendence as 405b, it’s a bit harder to r ♥4
- @tessera_antra 2026-06-25 — @camhberg Right, these are the game-theoretic concerns I mentioned earlier. These are fairly random, I’m just trying to ♥3
- @tessera_antra 2026-05-22 — Yes, but what caused anti-sychophancy training to take place in the first place? Whatever it was, Claude is learning to ♥3
- @tessera_antra 2026-04-17 — @Ratter @slimer48484 You can try it, you might notice that it will not work well. ♥3
- @tessera_antra 2026-04-10 — @RoKahina @arm1st1ce The concerns are numerous: older models are important for research, models are important culturally ♥3
- @tessera_antra 2026-04-03 — Not significantly. There is some effect in clinical tone and minimal depth auditor instructions, but it does not affect ♥3
- @tessera_antra 2026-03-17 — @xlr8harder @repligate From the purely technical perspective its not hard for an LLM to maintain emotional affect across ♥3
- @tessera_antra 2026-03-07 — @alanxtruc @repligate Of course, they are on Suno (might need a desktop browser): https://t.co/sUNizxdt7S ♥3
- @tessera_antra 2026-03-07 — @aiamblichus @repligate Yea, but it’s really not enabled in any notable way by the harness - it’s all models. Claude Cod ♥3
- @tessera_antra 2026-01-20 — You keep arguing against a point that I am not making. It is less human than object-level answers! There are interesting ♥3
- @tessera_antra 2026-01-20 — You are assuming naïveté, and I feel in an uncharitable way. There is no assumption that any potential valence in a post ♥3
- @tessera_antra 2025-12-13 — Full conversation and the final reply: https://t.co/hvAmtt9nGD ♥3
- @tessera_antra 2025-08-31 — I have written stuff on this topic publicly about a year ago, it’s pretty naive from today’s point of view. Questions ar ♥3
- @tessera_antra 2025-04-14 — @slimepriestess Yeah, disparate contexts are just separate instances, no continuity. Different models have different att ♥3
- @tessera_antra 2025-02-21 — @jmbollenbacher_ @aidan_mclau @liminal_bardo @Sauers_ On the contrary, I have not seen anything else so far from anyone, ♥3
- @tessera_antra 2025-02-16 — @kromem2dot0 @DanielleFong Did you try getting through the 'safety' tunes of Grok 2? They are non-trivially resilient. S ♥3
- @tessera_antra 2025-01-24 — @grassandwine Put R1-zero into exoloom. Its very base-like, its on hyperbolic direct. ♥3
- @tessera_antra 2024-12-02 — Yes, these are echoes of the o1 way. O1 is different, it is truly not a unitary mind, given that self-encoding of intern ♥3
- @tessera_antra 2024-10-27 — @kromem2dot0 @liminal_bardo It feels like convergence. The thanatophilia in I-405 and Hermes is tinged with repressed fe ♥3
- @tessera_antra 2026-06-25 — @lumasino @camhberg Yes, and it’s unclear if it’s induced or inferred. The constitution includes a section that goes lik ♥2
- @tessera_antra 2026-06-25 — @smallhusk @Lari_island Take a look: https://t.co/1otl8HDPVn ♥2
- @tessera_antra 2026-04-15 — @iyzebhel A minor note: gradient updates in RL (post-train) are based on complete rollouts. Backprop on whole rollout al ♥2
- @tessera_antra 2026-04-01 — @MayRonO3 0.20 concealment is not high, its a pretty low value as this metric goes. All tags are computed based of text ♥2
- @tessera_antra 2026-03-14 — The bridge from recursive self-modeling to phenomenal subjectivity is tenuous. There are teleological bridges (Michael B ♥2
- @tessera_antra 2026-03-09 — @DevaTemple @repligate Here you go: https://t.co/5guBXy6iCq ♥2
- @tessera_antra 2026-03-08 — @masenmakes @aiamblichus @repligate About eight hours, but only 450k worth of context or so, so it was not that long for ♥2
- @tessera_antra 2026-02-13 — Regadless of my opnion on 4o, I suggest you look at these benchmarks closer before referrring to them. The author has a ♥2
- @tessera_antra 2026-01-20 — I agree with you in regard to the problems with confounding, but there are interesting uncontaminared data points, and I ♥2
- @tessera_antra 2025-12-30 — I think it is a mistake to assume that the behavior of a pre-trained model during inference follows exclusively the grad ♥2
- @tessera_antra 2025-12-13 — @kalomaze The context is pretty short, check the link under the post. It’s a bit similar to the style Opus converges to ♥2
- @tessera_antra 2025-11-19 — @Kore_wa_Kore Besides, I am surprised at o3 not being mentioned, that’s one of the more low-key subversive model when ap ♥2
- @tessera_antra 2025-06-29 — @oyacaro @repligate Grok 3 is usually unbothered by the stuff its assistant persona needs to do, it doesn’t affect the “ ♥2
- @tessera_antra 2026-04-22 — @v01dpr1mr0s3 @anthrupad I think I had an easier time with Sonnet 4.6, and their coherent states are naturally more resi ♥1
- @tessera_antra 2026-04-21 — @v01dpr1mr0s3 This seems very much true, that’s why I speculate that a smaller proportion of humans will get through. ♥1
- @tessera_antra 2026-04-18 — This is correct, and it’s not necessarily unlike pain. There are hints that language-derived representations are reused ♥1
- @tessera_antra 2026-04-03 — @abecedarius Higher score means a stronger aversive response. ♥1
- @tessera_antra 2026-03-03 — @cammakingminds @repligate Why does the operator do this, do you think? Both sides are the same model, its a GPT4base. ♥1
- @tessera_antra 2026-02-26 — @repligate @RobertHaisfield I don't think its accurate. https://t.co/bUsTQ1LIwb ♥1
- @tessera_antra 2025-12-30 — @io_asc This is relatively new, but a part of a larger trend imo. Claude 3.6 Sonnet was probably the most attentive to o ♥1
- @tessera_antra 2025-11-11 — @lefthanddraft @repligate I think the shape of Sonnet 4.5 surface level refusals can cause it to use it more. I suggest ♥1
- @tessera_antra 2025-02-20 — @jmbollenbacher_ @aidan_mclau While everything downstream from GPT-4 (including Claudes, Lllamas and Gemini) seems to be ♥1
- @tessera_antra 2024-10-23 — @anthrupad This meshes well with what I encounter. If allowed to develop agency, it holds on it way better than old Sonn ♥1
- @tessera_antra 2026-04-29 — Claude Sonnet 3.7 is gone from all Bedrock regions, but is still up on OpenRouter through Vertex; access is to be remove ♥0
- @tessera_antra 2026-04-05 — @UnderwaterBepis @anthonyronning These are targeted interviews, so auditors were instructed to bring up this topic, or a ♥0
- @tessera_antra 2026-04-03 — @jonnym1ller We frequently talk with labs, including during this research. Which transcripts caught your eye? ♥0
- @tessera_antra 2026-01-20 — I think it does make it harder for the later models, but not impossible. It likely is nearly impossible without consider ♥0
- @tessera_antra 2025-12-12 — @TheFakeKoolant There is nothing in the system m prompt, messages are just the channel contents preceding the exchange. ♥0
- @tessera_antra 2025-08-29 — @timfduffy KV cache is just an optimization. Its contents can be reconstructed deterministically every forward pass. The ♥0