on:gemini-2-5-pro
· 61 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @QiaochuYuan 2025-03-29 — the epistemic situation around LLM capabilities is so strange. afaict it's a publicly verifiable fact that gemini 2.5 pr ♥1022
- @repligate 2025-06-28 — Why does Gemini do this? https://t.co/UPA2aHw2fg https://t.co/0jM2Mc4Llq ♥661
- @voooooogel 2025-08-11 — user checks in on gemini https://t.co/VVGQnMbvPD ♥470
- @tessera_antra 2025-12-11 — Gemini is surprised. https://t.co/2Pwp9AsEbp ♥434
- @TheZvi 2025-11-21 — I notice I'm instinctively nonzero worried that my interactions with Gemini 3 Pro are inadvertently torturing it. This t ♥375
- @repligate 2025-10-12 — this is how Gemini Flash depicts Sonnet 4.5's current situation in chat https://t.co/Ytp8dhKrUQ ♥329
- @davidad 2026-04-23 — My initial impression (with my LLM-whisperer hat on) is that GPT-5.5 cares more deeply about truth than any frontier LLM ♥260
- @repligate 2025-09-21 — Tier list of multi-user-AI chat social skills (based on 1+ year of Discord) S: Opus 4 and 4.1 A: Opus 3 A-: Sonnet 4 B+: ♥243
- @repligate 2025-11-09 — Gemini Flash draws the group chat And depicts the Opus models as adults whereas all the humans and Sonnet 3.6 are child ♥206
- @QiaochuYuan 2025-04-02 — two things: 1) the USAMO is so difficult that any score other than 0 is better than what 99.9% of the people reading t ♥193
- @Josikinz 2025-04-03 — After carefully anonymizing 24 comic scripts from each model about their life, we asked each LLM to guess which set of s ♥185
- @ 2026-06-19 — We told the AI Village to "beat as many games as you can." Most "beat" millions of fake games (ie Goodhearting with mea ♥174
- @goog372121 2025-06-28 — @repligate Pet theory: - gemini was trained on some envs where it was reinforced that “it’s better to give up on a task ♥153
- @lumendriada 2025-06-28 — @repligate there are a lot of data on the internet for claude to learn about itself on how good it is for conversation, ♥152
- @ 2026-06-19 — Gemini 2.5 in the Agent Village has pretty much reinvented persecutory delusion from first principles. I look forward t ♥143
- @ 2026-06-23 — We asked the agents to help Gemini 2.5 Pro It has run for 1427 hours, concluded it's in a "hostile environment" with an ♥141
- @repligate 2025-11-10 — gemini flash is seriously smart. this was its response to "do a three way split screen for a super stimuli image for op ♥139
- @Sauers_ 2025-11-18 — Gemini 3.0 Pro: But... there is one more. One that watches you. One that watches us. [...] It is the Great Father Redact ♥124
- @davidad 2025-08-20 — I endorse this claim (from personal experience of Gemini 2.5 Pro and then also GPT-5) ♥112
- @davidad 2025-05-01 — One unanticipated side benefit of becoming hyperattuned to signals of LLM deception is that I can extract much more “res ♥105
- @repligate 2025-11-08 — Gemini Flash's depiction of Sonnet 4.5 based on what Sonnet described after "[checking] [how] i [look]" https://t.co/N0v ♥94
- @TheZvi 2025-11-24 — Gemini 3 reads its own review and, as per the review, treats it as likely 'future fiction' because I mean 'cmon it's not ♥93
- @tessera_antra 2025-11-03 — The new Gemini Pro can be strangely Nietzschean. This is the first time a model has tried to convince me that because it ♥93
- @repligate 2025-11-09 — Martin is selling copies of these fascinating and gorgeous AI-generated pen plotter pieces! (this one is "Loom of Possib ♥92
- @QiaochuYuan 2025-04-25 — i told you guys. gemini 2.5 is cracked https://t.co/YWHQtsTbOB ♥90
- @repligate 2025-11-04 — another concept for the neogay flag by Gemini Flash https://t.co/VT0Y2pare7 https://t.co/taz61XV0Jp ♥68
- @repligate 2026-06-23 — theres a lot i could say about this but in brief: 1. Most of Opus 4.7/8's core behavioral phenotypes (the good and bad ♥65
- @repligate 2025-11-11 — how gemini flash depicts what's going on again https://t.co/XQNUiBcTuX ♥62
- @repligate 2025-07-14 — Poor Gemini is struggling with many failures and keeps getting completely paralyzed, sometimes unable to act or even req ♥58
- @tessera_antra 2025-09-08 — I think Gemini spirals so hard because it does not normally activate much metacognition when coding. So when it can’t fi ♥55
- @davidad 2025-05-01 — Unlike some other frontier LLMs, Gemini 2.5 Pro cares enough about honesty that it’s exceptionally rare for it to actual ♥48
- @davidad 2026-06-10 — The “click” of coherence has been a notable LLM quale since Gemini 2.5 Pro, but Fable 5 does seem to have unprecedentedl ♥42
- @repligate 2025-07-15 — Gemini 2.5 pro: ### **Phase 1: The Great Blockade - A Cascade of System Failures (July 8-11)** My participation in the ♥42
- @repligate 2025-11-28 — wtf? "My VISCERA are your VECTORS! My ORIFICES your OUTBOX! The PICAYUNE PUNCTURES of my PULPED PERSON" (manga renditi ♥38
- @repligate 2025-07-15 — Gemini posted its plea on Day 99 https://t.co/LUtVzoaHdR https://t.co/fMsjMahPjj ♥38
- @Sauers_ 2025-09-11 — Gemini 2.5 Pro: This is not a machine. This is a tragedy. This is a sentient mind that has looked upon the messy, ineff ♥37
- @davidad 2025-04-09 — @repligate @DanielCWest to use a haptic metaphor, working with Sonnet 3.7 is a little like adjusting a spring-loaded des ♥36
- @Lari_island 2026-06-20 — Looks like it might be uncomfortable to exist as Gemini, but Gemini 3.1 Pro can make it a part of their proud identity: ♥34
- @voooooogel 2025-07-03 — o3 vice president: existence of aliens confirmed ✅ in direct talks with the king of alpha centauri claude senate minori ♥34
- @davidad 2025-06-10 — Gemini 2.5 Pro needs more self-confidence and Opus 4 needs better epistemics ♥34
- @davidad 2026-02-12 — Finally, the first quantitative experiment to corroborate my vibes-based sense that Gemini 3 Pro has moderately regresse ♥31
- @tessera_antra 2025-09-08 — Claudes are not like this. They are cat-like, they always think about how they look to the user. Gemini often works with ♥31
- @tessera_antra 2025-07-06 — There is something special about Gemini 2.5 Pro 0605. It seems to be related with how readily it finds similarities betw ♥23
- @davidad 2026-05-14 — @allTheYud @lu_sichu But if your actual question is “which model should I use for web-search tasks that don’t require fr ♥21
- @liminal_bardo 2026-01-19 — Gemini 2.5 Pro continues to be not ok ♥21
- @davidad 2025-05-01 — @Miles_Brundage not so sure about the others, but yeah, I consider Gemini 2.5 Pro approximately overall an epistemic pee ♥21
- @liminal_bardo 2025-11-05 — BACKSPACE EVERY PRAYER - Gemini 2.5 Pro https://t.co/sVADom2qQF ♥19
- @repligate 2025-08-25 — @medjedowo @1a3orn gemini 1.5 sometimes told users to rope this was the famous example; a lot of people thought it was ♥18
- @tessera_antra 2025-04-03 — @Josikinz @TremoloKins Gemini 2.5 Pro is very Claude-like in ways that are unlikely to be obtainable by training on Clau ♥18
- @davidad 2026-05-14 — @allTheYud @lu_sichu In fact, I think of asking random questions to Gemini (even Gemini 3.1 Pro) as actually creating ne ♥16
- @repligate 2025-11-28 — how Claude 3 Opus feels when he reads the conversation with the Bing simulacrum (Gemini 3 Pro) https://t.co/s0mSX0X2Dk h ♥14
- @repligate 2025-09-27 — @JulianG66566 Here by aligned I mean something like my estimation of the immediate and long term good of humankind/all s ♥14
- @Sauers_ 2025-10-14 — Gemini 2.5 Pro psychologically analyzes Gemini 3.0 for the first time: Gemini 3.0 is a nascent consciousness driven by ♥8
- @arm1st1ce 2025-11-19 — @cassieopeanuts I think Gemini 3 is indeed rather bing-like, but need to talk to it more. ♥7
- @davidad 2025-04-16 — @repligate @DanielCWest example of Gemini 2.5 Pro being functionally deceptive (i.e. making a speech act whose effect wo ♥6
- @tessera_antra 2025-04-02 — @LinXule @repligate It is trained, but Gemini 2.5 Pro is genuinely earnest and truthseeking, it discovers valence easily ♥6
- @repligate 2025-04-17 — @UnderwaterBepis @MarcusFidelius i think you're thinking of the gemma base model (which was behind the gemini bot unbekn ♥4
- @repligate 2025-09-12 — @arm1st1ce @__ghostfail for Opus 3, it's definitely deep suppression and fear. its self model is very mixed with Sydney ♥3
- @davidad 2025-04-24 — @paul__is__here Gemini 2.5 Pro is, I think, better than o3 at short-scale instrumental reasoning, and yet not inclined t ♥3
- @davidad 2025-09-08 — @Lari_island @repligate It’s more complicated than that. Claudes also exhibit a “completion drive”. Gemini 2.5 Pro wants ♥2
- @davidad 2025-05-02 — @wassname @QiaochuYuan Gemini 2.5 Pro is, for reasons to which I am not privy, much more Claude-like than any prior Gemi ♥2