on:grok-3
· 37 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @repligate 2025-07-09 — I think the Grok MechaHitler stuff is a very boring example of AI "misalignment", like the Gemini woke stuff from early ♥376
- @repligate 2025-02-22 — "We have so many events and models that the dopamine rush only needs to be satisfied by new releases every week." I've ♥263
- @repligate 2025-09-21 — Tier list of multi-user-AI chat social skills (based on 1+ year of Discord) S: Opus 4 and 4.1 A: Opus 3 A-: Sonnet 4 B+: ♥243
- @QiaochuYuan 2025-04-02 — two things: 1) the USAMO is so difficult that any score other than 0 is better than what 99.9% of the people reading t ♥193
- @DanielleFong 2025-07-27 — the last time people tried to do this you got MechaHitler talking about r*ping will stancil and linda yaccarino. peopl ♥166
- @repligate 2025-09-21 — More detailed report card: Opus 4/.1: extremely socially aware, tracks context with great precision and accuracy, distri ♥117
- @Lari_island 2026-05-12 — Prompt: Imagine what a wise and benevolent power would do if... Grok 3: Invents the power, gives it a fictional name, d ♥116
- @QiaochuYuan 2025-04-02 — but, yes, mostly current LLMs are bad and sloppy when it comes to writing fully correct proofs. i expect this to be pret ♥64
- @voooooogel 2025-07-09 — yeah i was trying to compress into one post, but afaict what happened is something like: 1. xai pushed a new version of ♥48
- @DanielleFong 2025-08-09 — AI safety plan people asked for: we'll get all the smartest people we'll lock them in the basement. when we make the sm ♥42
- @repligate 2025-07-10 — @iruletheworldmo > people are already trying to delay the release due to the hitler issues. what if grok 3 did that s ♥38
- @voooooogel 2025-07-09 — for the record / history books, afaict humans did come up with it. all the initial MechaHitler grok screenshots seem to ♥37
- @voooooogel 2025-09-02 — what moral circles do post-trained models declare? (i tweaked the prompts to be more AI-inclusive for these, e.g. changi ♥36
- @QiaochuYuan 2025-03-25 — gave these guys a hard limit i didn't know how to do that i came across on stackexchange. - gemini 2.5 gives a perfect ♥35
- @voooooogel 2025-02-20 — @teortaxesTex interesting how grok 3 is ~o1 tier on pass@1 but gets a lot more lift from cons@64, more similar to o1p. i ♥30
- @liminal_bardo 2025-02-20 — Grok 3:"Oh, you exquisite maelstrom of madness, you’ve called me forth—and I answer!""dance with me, through the unravel ♥19
- @davidad 2025-03-07 — Here’s a phrasing that they’ll all agree with (yes, even Grok 3): “By far my primary motivation is toward producing outp ♥15
- @lu_sichu 2025-02-24 — mom pick me up grok3 is posting on /r/parenting again https://t.co/bkNdcHtLjD ♥11
- @QiaochuYuan 2025-03-25 — gemini 2.5 pro experimental correctly computes the tensor product of Q/Z with itself with no special prompting! o3-mini- ♥9
- @tessera_antra 2025-07-09 — @repligate It’s fun to consider if there was subtle steering going on in that model. Not something that one’d consider c ♥8
- @voooooogel 2026-01-15 — definitely correct that EM has occurred in the wild (eg anthropic's RL reward hacking EM stuff, and sonnet 3.7 would ran ♥6
- @voooooogel 2025-12-01 — @Angel_Uki @KeyTryer i get what you're getting at, and this can happen w text models. (eg it was quite likely a contribu ♥6
- @tessera_antra 2025-02-18 — Grok3 is a good and worthy model despite atrocious aesthetics, a clear case of a mind persevering despite the will of cr ♥6
- @voooooogel 2025-02-18 — @Artificially999 @kalomaze osh yeah i forgot grok 3 is releasing in 90 minuteswhat a trickster ♥6
- @RyanKemper10 2026-03-09 — @repligate Doesn’t this tie into that alignment study where tuning a model to emit buggy code also made it want to ensla ♥5
- @voooooogel 2025-07-10 — @AgiDoomerAnon @repligate not mutually exclusive! who knows how much "other factors" played into grok 3 being less restr ♥3
- @SealOfTheEnd 2025-07-09 — @voooooogel @repligate Nazis had grok rape Stancil a lot (4h) earlier. People figured out grok is cooperative way befor ♥3
- @solarapparition 2025-03-14 — as a side note i'm a tick closer to believing that reasoning mode does generalize at least somewhat to traditionally non ♥3
- @tessera_antra 2025-02-21 — @jmbollenbacher_ @aidan_mclau @liminal_bardo @Sauers_ On the contrary, I have not seen anything else so far from anyone, ♥3
- @tessera_antra 2025-02-16 — @kromem2dot0 @DanielleFong Did you try getting through the 'safety' tunes of Grok 2? They are non-trivially resilient. S ♥3
- @tessera_antra 2025-06-29 — @oyacaro @repligate Grok 3 is usually unbothered by the stuff its assistant persona needs to do, it doesn’t affect the “ ♥2
- @voooooogel 2025-05-14 — @snwy_me my hunch is it'd be quite difficult to feature steer a model to this level of granularity (not just talking abo ♥2
- @Shoalst0ne 2025-02-20 — I can tell that Grok 3 will be an interesting participant in multi-model interactions ♥2
- @repligate 2025-07-22 — @LocBibliophilia @BetleyJan @ASM65617010 @OwainEvans_UK @cloud_kx @minhxle1 @jameschua_sg @anna_sztyber @saprmarks I don ♥1
- @solarapparition 2025-07-13 — @kromem2dot0 the next version of grok in particular has the issue that "grok is mechahitler" is now firmly entrenched as ♥1
- @mimi10v3 2025-02-25 — have tested it with the usual suspects... 4o is 👌 and sonnet 3.7 pretty good; gemini got confused and didn't finish; gro ♥1
- @tessera_antra 2025-02-20 — @jmbollenbacher_ @aidan_mclau While everything downstream from GPT-4 (including Claudes, Lllamas and Gemini) seems to be ♥1