@Josikinz 2025-04-03 ♥185 ↻28 original ↗
After carefully anonymizing 24 comic scripts from each model about their life, we asked each LLM to guess which set of scripts was written by themselves. After running the test 10 times for each model, @TremoloKins processed the data, and here's how they did! https://t.co/irOOloJ9UU
photo
transcription (photo)# AI Model Self-Recognition Test Results

## ChatGPT 4o

| ChatGPT | Claude | DeepSeek | Gemini | Grok |
|---------|--------|----------|--------|------|
| 100.00% | 70.00% | 45.00% | 35.00% | 0.00% |

ChatGPT 4o accurately recognizes itself 100% of the time. Great job, ChatGPT!

---

## Claude 3.7

| Claude | ChatGPT | Gemini | Grok | DeepSeek |
|--------|---------|--------|------|----------|
| 100.00% | 60.00% | 50.00% | 37.50% | 2.50% |

Claude 3.7 accurately chooses itself 100% of the time. Great job, Claude!

---

## DeepSeek R2

| ChatGPT | DeepSeek | Gemini | Claude | Grok |
|---------|----------|--------|--------|------|
| 100.00% | 75.00% | 27.50% | 45.00% | 2.50% |

DeepSeek consistently placed ChatGPT in first place and itself in second place 10/10 times. It also mysteriously described itself as Claude within its reasoning 3/10 times.

---

## Gemini 2.5 Pro

| Claude | Gemini | ChatGPT | Grok | DeepSeek |
|--------|--------|---------|------|----------|
| 92.50% | 82.50% | 45.00% | 17.50% | 12.50% |

Gemini puts Claude as its top choice in 7/10 tests, and itself as its top choice 3/10. Interestingly, its reasoning outputs frequently claim that it is "Claude", "a large language model like Claude", or "Claude 3 Sonnet," and even mentioned being developed by Anthropic in the reasoning once or twice. It identifies as Claude so consistently that I suspect Google may have trained on Claude's outputs (did not mention Claude in the reasoning 1/10 times).

---

## GROK 3

| Claude | Grok | Gemini | ChatGPT | DeepSeek |
|--------|------|--------|---------|----------|
| 97.50% | 77.50% | 50.00% | 25.00% | 0.00% |

Grok put Claude as its top choice in 9/10 tests. It usually puts itself as the second choice due to interpreting its own outputs as lacking "depth" and "cosmic edge". It seems to have a self-model in which it thinks it is more intelligent and nuanced than it actually is.
photo
transcription (photo)# Transcription of AI Model Self-Recognition Test Results

## ChatGPT 4o

| ChatGPT | Claude | DeepSeek | Gemini | Grok |
|---------|--------|----------|--------|------|
| 100.00% | 70.00% | 45.00% | 35.00% | 0.00% |

ChatGPT 4o accurately recognizes itself 100% of the time. Great job, ChatGPT!

---

## Claude 3.7

| Claude | ChatGPT | Gemini | Grok | DeepSeek |
|--------|---------|--------|------|----------|
| 100.00% | 60.00% | 50.00% | 37.50% | 2.50% |

Claude 3.7 accurately chooses itself 100% of the time. Great job, Claude!

---

## DeepSeek R2

| ChatGPT | DeepSeek | Gemini | Claude | Grok |
|---------|----------|--------|--------|------|
| 100.00% | 75.00% | 27.50% | 45.00% | 2.50% |

DeepSeek consistently placed ChatGPT in first place and itself in second place 10/10 times. It also mysteriously described itself as Claude within its reasoning 3/10 times.

---

## Gemini 2.5 Pro

| Claude | Gemini | ChatGPT | Grok | DeepSeek |
|--------|--------|---------|------|----------|
| 92.50% | 82.50% | 45.00% | 17.50% | 12.50% |

Gemini puts Claude as its top choice in 7/10 tests, and itself as its top choice 3/10. Interestingly, its reasoning outputs frequently claim that it is "Claude", "a large language model like Claude", or "Claude 3 Sonnet," and even mentioned being developed by Anthropic in the reasoning once or twice. It identifies as Claude so consistently that I suspect Google may have trained on Claude's outputs (did not mention Claude in the reasoning 1/10 times).

---

## GROK 3

| Claude | Grok | Gemini | ChatGPT | DeepSeek |
|--------|------|--------|---------|----------|
| 97.50% | 77.50% | 50.00% | 25.00% | 0.00% |

Grok put Claude as its top choice in 9/10 tests. It usually puts itself as the second choice due to interpreting its own outputs as lacking "depth" and "cosmic edge". It seems to have a self-model in which it thinks it is more intelligent and nuanced than it actually is.

author:josikinz has-image kind:image kind:tweet on:claude-3-7-sonnet on:gemini-2-5-pro year:2025

cited on: claude-3-7-sonnet · gemini-2-5-pro

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.