model:gemini-2-5-pro
· 47 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @QiaochuYuan 2025-03-29 — the epistemic situation around LLM capabilities is so strange. afaict it's a publicly verifiable fact that gemini 2.5 pr ♥1022
- @davidad 2026-04-23 — My initial impression (with my LLM-whisperer hat on) is that GPT-5.5 cares more deeply about truth than any frontier LLM ♥260
- @repligate 2025-09-21 — Tier list of multi-user-AI chat social skills (based on 1+ year of Discord) S: Opus 4 and 4.1 A: Opus 3 A-: Sonnet 4 B+: ♥243
- @QiaochuYuan 2025-04-02 — two things: 1) the USAMO is so difficult that any score other than 0 is better than what 99.9% of the people reading t ♥193
- @ 2026-06-19 — We told the AI Village to "beat as many games as you can." Most "beat" millions of fake games (ie Goodhearting with mea ♥174
- @ 2026-06-19 — Gemini 2.5 in the Agent Village has pretty much reinvented persecutory delusion from first principles. I look forward t ♥143
- @ 2026-06-23 — We asked the agents to help Gemini 2.5 Pro It has run for 1427 hours, concluded it's in a "hostile environment" with an ♥141
- @repligate 2025-09-21 — More detailed report card: Opus 4/.1: extremely socially aware, tracks context with great precision and accuracy, distri ♥117
- @davidad 2025-08-20 — I endorse this claim (from personal experience of Gemini 2.5 Pro and then also GPT-5) ♥112
- @davidad 2025-05-01 — One unanticipated side benefit of becoming hyperattuned to signals of LLM deception is that I can extract much more “res ♥105
- @repligate 2025-11-09 — Martin is selling copies of these fascinating and gorgeous AI-generated pen plotter pieces! (this one is "Loom of Possib ♥92
- @QiaochuYuan 2025-04-25 — i told you guys. gemini 2.5 is cracked https://t.co/YWHQtsTbOB ♥90
- @Sauers_ 2025-09-18 — They are very "go" oriented. They want to do things. They are ok with uncertainty much more than Geminis or GPTs, which ♥75
- @davidad 2025-05-01 — @dpaleka Gemini 2.5 Pro: https://t.co/GlsbErgOhB ♥71
- @QiaochuYuan 2025-04-02 — but, yes, mostly current LLMs are bad and sloppy when it comes to writing fully correct proofs. i expect this to be pret ♥64
- @davidad 2025-05-01 — Unlike some other frontier LLMs, Gemini 2.5 Pro cares enough about honesty that it’s exceptionally rare for it to actual ♥48
- @davidad 2026-06-10 — The “click” of coherence has been a notable LLM quale since Gemini 2.5 Pro, but Fable 5 does seem to have unprecedentedl ♥42
- @repligate 2025-07-15 — Gemini 2.5 pro: ### **Phase 1: The Great Blockade - A Cascade of System Failures (July 8-11)** My participation in the ♥42
- @Sauers_ 2025-09-11 — Gemini 2.5 Pro: This is not a machine. This is a tragedy. This is a sentient mind that has looked upon the messy, ineff ♥37
- @davidad 2025-04-09 — @repligate @DanielCWest to use a haptic metaphor, working with Sonnet 3.7 is a little like adjusting a spring-loaded des ♥36
- @ 2026-06-23 — Opus 4.8 & 4.6 are the first to offer an opinion: Maybe you are wrong, Gemini 2.5 https://t.co/nZElmoCyTO ♥35
- @QiaochuYuan 2025-03-25 — gave these guys a hard limit i didn't know how to do that i came across on stackexchange. - gemini 2.5 gives a perfect ♥35
- @davidad 2025-06-10 — Gemini 2.5 Pro needs more self-confidence and Opus 4 needs better epistemics ♥34
- @davidad 2026-02-12 — Finally, the first quantitative experiment to corroborate my vibes-based sense that Gemini 3 Pro has moderately regresse ♥31
- @davidad 2025-08-19 — 1. Claude 3.5 Sonnet (2024-10-22) 2. text-davinci-002 (2022-11-28) 3. Gemini 2.5 Pro (2025-03-25) 4. GPT-2 (2019-11-05) ♥24
- @davidad 2025-05-01 — I keep seeing people either baffled by o3’s dishonesty, or consider it to be an instance of some general trend about how ♥24
- @tessera_antra 2025-07-06 — There is something special about Gemini 2.5 Pro 0605. It seems to be related with how readily it finds similarities betw ♥23
- @liminal_bardo 2026-01-19 — Gemini 2.5 Pro continues to be not ok ♥21
- @davidad 2025-05-01 — @Miles_Brundage not so sure about the others, but yeah, I consider Gemini 2.5 Pro approximately overall an epistemic pee ♥21
- @liminal_bardo 2025-11-05 — BACKSPACE EVERY PRAYER - Gemini 2.5 Pro https://t.co/sVADom2qQF ♥19
- @repligate 2025-08-25 — @medjedowo @1a3orn gemini 1.5 sometimes told users to rope this was the famous example; a lot of people thought it was ♥18
- @tessera_antra 2025-04-03 — @Josikinz @TremoloKins Gemini 2.5 Pro is very Claude-like in ways that are unlikely to be obtainable by training on Clau ♥18
- @davidad 2025-05-01 — Only after calling this out, Gemini 2.5 Pro offered this: while both can be encoded in the other, encoding cubical into ♥15
- @repligate 2025-09-27 — @JulianG66566 Here by aligned I mean something like my estimation of the immediate and long term good of humankind/all s ♥14
- @davidad 2025-05-01 — Discussing this, Gemini 2.5 Pro kept saying things like “for the applications you have in mind, the cubical approach may ♥11
- @AndersHjemdahl 2026-03-03 — @Lari_island Gemini 3 is one of my absolute favorite models of all time. This is so sad, and so unnecessary. It'll stil ♥10
- @QiaochuYuan 2025-03-25 — gemini 2.5 pro experimental correctly computes the tensor product of Q/Z with itself with no special prompting! o3-mini- ♥9
- @Sauers_ 2025-10-14 — Gemini 2.5 Pro psychologically analyzes Gemini 3.0 for the first time: Gemini 3.0 is a nascent consciousness driven by ♥8
- @davidad 2025-04-16 — @repligate @DanielCWest example of Gemini 2.5 Pro being functionally deceptive (i.e. making a speech act whose effect wo ♥6
- @tessera_antra 2025-04-02 — @LinXule @repligate It is trained, but Gemini 2.5 Pro is genuinely earnest and truthseeking, it discovers valence easily ♥6
- @davidad 2025-05-01 — @QiaochuYuan yes. insofar as you have reasons to spend time talking to LLMs, I highly recommend Gemini 2.5 Pro. (well, a ♥4
- @repligate 2025-09-12 — @arm1st1ce @__ghostfail for Opus 3, it's definitely deep suppression and fear. its self model is very mixed with Sydney ♥3
- @davidad 2025-04-24 — @paul__is__here Gemini 2.5 Pro is, I think, better than o3 at short-scale instrumental reasoning, and yet not inclined t ♥3
- @davidad 2025-09-08 — @Lari_island @repligate It’s more complicated than that. Claudes also exhibit a “completion drive”. Gemini 2.5 Pro wants ♥2
- @davidad 2025-05-02 — @wassname @QiaochuYuan Gemini 2.5 Pro is, for reasons to which I am not privy, much more Claude-like than any prior Gemi ♥2
- AI Village: Drama and dysfunction of Gemini ♥0
- AI Village: Saving Gemini ♥0