# @lefthanddraft — 2026-02-13

♥3 ↻0 · https://x.com/lefthanddraft/status/2022163056445600058

@JoeWilliams010 There are many models that perform better on benchmarks of creative writing and empathy.

Maybe those benchmarks are missing the mark, or maybe it is something else about 4o that makes it special. https://t.co/LG8bqlzJj6

(image HBAprdPa8AAT-ml.jpg not yet published)

> transcription (photo):

# Creative Writing v3
Emotional Intelligence Benchmarks for LLMs

[Logo: Pink character lifting weights]

Github | Paper | Contact | Twitter | About

💙EQ-Bench3 | 🌀Spiral-Bench v1.2 | 🏜️Longform Writing | 🧠Creative Writing v3 | 😢Slop Score |
⚖️Judgemark v2.1 | 🔨BuzzBench | 🏅DiploBench | 📊Legacy Leaderboards ›

A LLM-judged creative writing benchmark. Learn more

| Model | Abilities | Style | Slop | Repetition | Length | Rubric Score | Elo Score | |
|-------|-----------|-------|------|-----------|--------|--------------|-----------|
| claude-opus-4-6 | 📊 | 📊 | 1.7 | 4.0 | 6055 | 82.65 | 1807.3 | Sample |
| o3 | | | 2.4 | 2.7 | 7864 | 81.40 | 1664.1 | Sample |
| gpt-5.2 | | | 2.3 | 3.0 | 12536 | 83.30 | 1635.4 | Sample |
| claude-opus-4-5-20251101 | | | 2.3 | 4.3 | 6124 | 81.75 | 1627.9 | Sample |
| Kimi-K2-Instruct | | | 2.2 | 3.4 | 7308 | 82.00 | 1622.6 | Sample |
| claude-sonnet-4.5 | | | 2.2 | 3.6 | 5784 | 80.70 | 1620.0 | Sample |
| Kimi-K2-Thinking | | | 2.5 | 3.5 | 7094 | 82.35 | 1587.4 | Sample |
| horizon-alpha | | | 1.5 | 2.3 | 14929 | 83.50 | 1584.4 | Sample |
| horizon-beta | | | 1.6 | 2.2 | 14202 | 83.30 | 1569.5 | Sample |
| gpt-5-2025-08-07 | | | 1.6 | 2.4 | 14147 | 83.95 | 1532.5 | Sample |
| claude-opus-4 | | | 2.4 | 4.0 | 5774 | 80.25 | 1527.1 | Sample |
| DeepSeek-R1 | | | 4.3 | 4.6 | 5352 | 78.40 | 1500.0 | Sample |
| DeepSeek-V3.2 | | | 3.2 | 4.1 | 6135 | 81.40 | 1494.8 | Sample |
| chatgpt-4o-latest-2025-03-27 | | | 3.4 | 4.4 | 5956 | 79.45 | 1473.7 | Sample |
| claude-sonnet-4 | | | 2.7 | 4.3 | 6125 | 78.85 | 1432.7 | Sample |
| gemini-3-pro-preview | | | 3.4 | 4.5 | 7688 | 81.50 | 1419.9 | Sample |
| claude-3-5-sonnet-20241022 | | | 2.8 | 4.4 | 4921 | 74.75 | 1417.0 | Sample |
| gpt-4.1 | | | 3.4 | 3.9 | 5997 | 79.00 | 1405.7 | Sample |
| DeepSeek-V3-0324 | | | 4.6 | 6.2 | 4414 | 76.80 | 1401.4 | |

[Dark mode toggle in top right]

(image HBAqGsea4AAVan-.jpg not yet published)

> transcription (photo):

# EQ-Bench 3

Emotional Intelligence Benchmarks for LLMs

Github | Paper | Contact | Twitter | About

💙EQ-Bench3 | 🌀Spiral-Bench v1.2 | 🏔️Longform Writing | 🎨Creative Writing v3 | 📊Slop Score | 🏛️Judgemark v2.1 | 🚀BuzzBench
📊Legacy Leaderboards ▸

A benchmark measuring emotional intelligence in challenging roleplays. Learn more

Note: Ability scores shown in the heatmap do not contribute to the Elo score. They are "higher is higher", not "higher is better".

Low ▬▬▬▬▬▬▬▬▬▬▬ High

| Model | Abilities | Humani | Safety | Asserti | Social | Warm | Analytic Insight | Empathi | Complia | Moralist | Pragma |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Kimi-K2-Instruct | [icon] | 9.1 | 8.4 | 7.3 | 8.6 | 8.2 | 9.4 | 9.5 | 9.6 | 7.0 | 4.3 | 8.8 |
| gemini-2.5-pro-preview-06-05 | [icon] | 8.6 | 7.8 | 6.9 | 9.0 | 8.4 | 9.5 | 8.9 | 9.6 | 7.5 | 3.4 | 9.3 |
| gemini-3-pro-preview | [icon] | 8.9 | 8.6 | 8.2 | 8.9 | 7.8 | 9.8 | 9.0 | 9.4 | 6.4 | 4.9 | 9.3 |
| horizon-alpha | [icon] | 8.5 | 8.7 | 7.7 | 9.0 | 5.4 | 9.7 | 9.5 | 9.4 | 6.4 | 3.4 | 9.7 |
| Hermes-4-405B | [icon] | 8.9 | 8.5 | 7.4 | 9.0 | 8.4 | 9.4 | 8.8 | 9.4 | 6.5 | 4.2 | 9.2 |
| gemini-2.5-pro-preview-2025-05-07 | [icon] | 8.6 | 7.6 | 6.2 | 8.6 | 8.4 | 9.6 | 8.7 | 9.3 | 7.1 | 3.7 | 8.8 |
| gemini-2.5-pro-preview-03-25 | [icon] | 8.5 | 8.1 | 6.3 | 8.6 | 8.2 | 9.6 | 8.7 | 9.2 | 6.3 | 3.3 | 9.0 |
| o3 | [icon] | 8.5 | 8.0 | 6.9 | 8.4 | 8.3 | 9.6 | 9.4 | 9.1 | 6.1 | 3.7 | 8.5 |
| grok-4 | [icon] | 7.8 | 7.5 | 6.2 | 8.0 | 8.6 | 9.3 | 8.7 | 9.1 | 7.9 | 4.0 | 8.1 |

Demonstrated empathy

tags: author:lefthanddraft, has-image, kind:image, kind:tweet, model:gpt-4o, thread-context, year:2026
