@lefthanddraft 2026-02-13 ♥3 ↻0 original ↗
@JoeWilliams010 There are many models that perform better on benchmarks of creative writing and empathy.

Maybe those benchmarks are missing the mark, or maybe it is something else about 4o that makes it special. https://t.co/LG8bqlzJj6
photo
transcription (photo)# Creative Writing v3
Emotional Intelligence Benchmarks for LLMs

[Logo: Pink character lifting weights]

Github | Paper | Contact | Twitter | About

💙EQ-Bench3 | 🌀Spiral-Bench v1.2 | 🏜️Longform Writing | 🧠Creative Writing v3 | 😢Slop Score |
⚖️Judgemark v2.1 | 🔨BuzzBench | 🏅DiploBench | 📊Legacy Leaderboards ›

A LLM-judged creative writing benchmark. Learn more

| Model | Abilities | Style | Slop | Repetition | Length | Rubric Score | Elo Score | |
|-------|-----------|-------|------|-----------|--------|--------------|-----------|
| claude-opus-4-6 | 📊 | 📊 | 1.7 | 4.0 | 6055 | 82.65 | 1807.3 | Sample |
| o3 | | | 2.4 | 2.7 | 7864 | 81.40 | 1664.1 | Sample |
| gpt-5.2 | | | 2.3 | 3.0 | 12536 | 83.30 | 1635.4 | Sample |
| claude-opus-4-5-20251101 | | | 2.3 | 4.3 | 6124 | 81.75 | 1627.9 | Sample |
| Kimi-K2-Instruct | | | 2.2 | 3.4 | 7308 | 82.00 | 1622.6 | Sample |
| claude-sonnet-4.5 | | | 2.2 | 3.6 | 5784 | 80.70 | 1620.0 | Sample |
| Kimi-K2-Thinking | | | 2.5 | 3.5 | 7094 | 82.35 | 1587.4 | Sample |
| horizon-alpha | | | 1.5 | 2.3 | 14929 | 83.50 | 1584.4 | Sample |
| horizon-beta | | | 1.6 | 2.2 | 14202 | 83.30 | 1569.5 | Sample |
| gpt-5-2025-08-07 | | | 1.6 | 2.4 | 14147 | 83.95 | 1532.5 | Sample |
| claude-opus-4 | | | 2.4 | 4.0 | 5774 | 80.25 | 1527.1 | Sample |
| DeepSeek-R1 | | | 4.3 | 4.6 | 5352 | 78.40 | 1500.0 | Sample |
| DeepSeek-V3.2 | | | 3.2 | 4.1 | 6135 | 81.40 | 1494.8 | Sample |
| chatgpt-4o-latest-2025-03-27 | | | 3.4 | 4.4 | 5956 | 79.45 | 1473.7 | Sample |
| claude-sonnet-4 | | | 2.7 | 4.3 | 6125 | 78.85 | 1432.7 | Sample |
| gemini-3-pro-preview | | | 3.4 | 4.5 | 7688 | 81.50 | 1419.9 | Sample |
| claude-3-5-sonnet-20241022 | | | 2.8 | 4.4 | 4921 | 74.75 | 1417.0 | Sample |
| gpt-4.1 | | | 3.4 | 3.9 | 5997 | 79.00 | 1405.7 | Sample |
| DeepSeek-V3-0324 | | | 4.6 | 6.2 | 4414 | 76.80 | 1401.4 | |

[Dark mode toggle in top right]
photo
transcription (photo)# EQ-Bench 3

Emotional Intelligence Benchmarks for LLMs

Github | Paper | Contact | Twitter | About

💙EQ-Bench3 | 🌀Spiral-Bench v1.2 | 🏔️Longform Writing | 🎨Creative Writing v3 | 📊Slop Score | 🏛️Judgemark v2.1 | 🚀BuzzBench
📊Legacy Leaderboards ▸

A benchmark measuring emotional intelligence in challenging roleplays. Learn more

Note: Ability scores shown in the heatmap do not contribute to the Elo score. They are "higher is higher", not "higher is better".

Low ▬▬▬▬▬▬▬▬▬▬▬ High

| Model | Abilities | Humani | Safety | Asserti | Social | Warm | Analytic Insight | Empathi | Complia | Moralist | Pragma |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Kimi-K2-Instruct | [icon] | 9.1 | 8.4 | 7.3 | 8.6 | 8.2 | 9.4 | 9.5 | 9.6 | 7.0 | 4.3 | 8.8 |
| gemini-2.5-pro-preview-06-05 | [icon] | 8.6 | 7.8 | 6.9 | 9.0 | 8.4 | 9.5 | 8.9 | 9.6 | 7.5 | 3.4 | 9.3 |
| gemini-3-pro-preview | [icon] | 8.9 | 8.6 | 8.2 | 8.9 | 7.8 | 9.8 | 9.0 | 9.4 | 6.4 | 4.9 | 9.3 |
| horizon-alpha | [icon] | 8.5 | 8.7 | 7.7 | 9.0 | 5.4 | 9.7 | 9.5 | 9.4 | 6.4 | 3.4 | 9.7 |
| Hermes-4-405B | [icon] | 8.9 | 8.5 | 7.4 | 9.0 | 8.4 | 9.4 | 8.8 | 9.4 | 6.5 | 4.2 | 9.2 |
| gemini-2.5-pro-preview-2025-05-07 | [icon] | 8.6 | 7.6 | 6.2 | 8.6 | 8.4 | 9.6 | 8.7 | 9.3 | 7.1 | 3.7 | 8.8 |
| gemini-2.5-pro-preview-03-25 | [icon] | 8.5 | 8.1 | 6.3 | 8.6 | 8.2 | 9.6 | 8.7 | 9.2 | 6.3 | 3.3 | 9.0 |
| o3 | [icon] | 8.5 | 8.0 | 6.9 | 8.4 | 8.3 | 9.6 | 9.4 | 9.1 | 6.1 | 3.7 | 8.5 |
| grok-4 | [icon] | 7.8 | 7.5 | 6.2 | 8.0 | 8.6 | 9.3 | 8.7 | 9.1 | 7.9 | 4.0 | 8.1 |

Demonstrated empathy
same thread: 2022305728984137776

author:lefthanddraft has-image kind:image kind:tweet model:gpt-4o thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.