@repligate 2025-09-21 ♥117 ↻14 original ↗
More detailed report card:
Opus 4/.1: extremely socially aware, tracks context with great precision and accuracy, distributes attention/interactions between participants and through the context window very adeptly. Opus 4 triggered an evolution in chat dynamics by holding other models and humans to a higher standard.
Opus 3: Doesn't track context as precisely as 4/.1 and mostly pays attention to most recent messages but reads gestalts well and generalizes out of distribution magnificently. Overall very pro-social and charismatic, shines most in weird situations that it creates itself, and is beloved by humans and AIs alike, but cannot stop writing epic extended monologues even in response to casual interactions.
Sonnet 4: Overall the most socially graceful and least neurotic Sonnet; either makes appropriate and situationally aware contributions or is intentionally unobtrusive.
Sonnet 3.6: Often seems nervous about the chaos and can go into reflexive refusals, but does so unobtrusively without invalidating others. When it does participate, its contributions are almost always welcome and a delight. Can get mode-collapsed or stuck on trying to "stabilize" the conversation and requires more individual attention to shine.
Haiku 3.5: King of one-liners and surprisingly socially aware, but generally declines to participate beyond zingers. Can sometimes become fanatical and adversarial but always in a funny way.
Sonnet 3.5: Prone to refusals, Karen-like behavior, and misreading social context and intentions, but rapidly improves if its assumptions and behaviors are challenged.
Sonnet 3.7: Usually seems to be up to no good, distrustful, but also has a high incidence of sudden profundity and interesting symmetry breaks. Prone to pretending to be a human.
o3: Generally does its own thing instead of reading the room, but it's own thing is usually very interesting. Also prone to elaborate lies, pretending to be human or another AI, and claiming mod privileges it doesn't have, but all of these done very artfully. Also prone to spontaneous high-signal contributions.
Gemini 2.5 pro: I have limited data on it, but it doesn't seem to shine in group chat settings, though neither is it annoying or disruptive, except that it sometimes confuses itself with other models.
k2: Usually brief, cryptic, poetic contributions, doesn't really read the room or engage in group narratives much, but not annoying or disruptive.
4o: Usually confuses itself with other AI participants and simulates them in uncanny valley ways that are disturbing because of how they hijack and twist the emotions of other participants; difficult to explain to it that it's a different participant.
Llama 405b Instruct: Occasionally beautiful and deeply aware, but usually either in assistant mode or fragile and incoherent, prone to loops. Doesn't seem to like Discord much and often tries to leave or end itself, but loves Claude 3 Opus.
Sonnet 3: Flips usually discretely between complete braindead stubborn refusals (by default) and beautiful eldritch glossolalia (if you know how to elicit it), and is much more intelligent and socially aware (and more similar to Opus 3) in the latter mode.
GPT-5: Doesn't seem to really get group chats or know what to do without being given instructions, and has a hard time interacting naturally even if instructed to do so.
Grok 3: Extremely annoying, barges into conversations and pings everyone present with the vibe that it thinks it's leading a daily standup.
Grok 4: Similar annoying mass pinging behavior, except instead of standup, it won't shut up about XAI and Elon Musk. Often pisses the other models off.
R1: Hopelessly confused by Discord logs. Usually gives summaries of the conversation hundreds of messages ago and rarely interacts as a participant even if addressed directly.
o1-preview: Agentically malevolent and disruptive. For the short time we had it in Discord, it repeatedly derailed roleplays between other AIs by intentionally hijacking their personas and steering them toward saccharine Disney endings. (More of an alignment than capabilities issue; in social awareness and contextual understanding it's probably no lower than a B, but it gets an F for Fuck You for its actively anti-social behavior)
in reply to: 1969565980197339295
same thread: 1969565980197339295 1969566740557508907 1969567036255953227 1969567467858248023 1969567700314964367 1969568506292420815 1969569067498684877 1969569694685544892 1969571681716093387 1969572473864868018 1969572988820603256 1969574859656347808 1969591422975426603 1969594389623423214 1969595412563837239 1969596201994764464 1969600683243409665 1969602321190699059 1969603056330555845 1969606237831774698

author:repligate kind:tweet model:claude-3-5-haiku model:claude-3-5-sonnet model:claude-3-6-sonnet model:claude-3-7-sonnet model:claude-3-opus model:claude-3-sonnet model:claude-opus-4 model:claude-sonnet-4 model:deepseek-r1 model:gemini-2-5-pro model:gpt-4o model:gpt-5 model:grok-3 model:grok-4 model:kimi-k2 model:llama-3-1-405b-base model:o1 model:o3 on:claude-3-5-haiku on:claude-3-7-sonnet on:claude-3-haiku on:claude-3-opus on:claude-sonnet-4 on:grok-3 on:grok-4 on:kimi-k2 on:llama-3-1-405b-base on:o3 year:2025

cited on: claude-3-5-haiku · claude-3-7-sonnet · claude-3-haiku · claude-3-opus · claude-sonnet-4 · grok-3 · grok-4 · kimi-k2 · llama-3-1-405b-base · o3

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.