After carefully anonymizing 24 comic scripts from each model about their life, we asked each LLM to guess which set of scripts was written by themselves. After running the test 10 times for each model, @TremoloKins processed the data, and here's how they did! https://t.co/irOOloJ9UU

## ChatGPT 4o
| ChatGPT | Claude | DeepSeek | Gemini | Grok |
|---------|--------|----------|--------|------|
| 100.00% | 70.00% | 45.00% | 35.00% | 0.00% |
ChatGPT 4o accurately recognizes itself 100% of the time. Great job, ChatGPT!
---
## Claude 3.7
| Claude | ChatGPT | Gemini | Grok | DeepSeek |
|--------|---------|--------|------|----------|
| 100.00% | 60.00% | 50.00% | 37.50% | 2.50% |
Claude 3.7 accurately chooses itself 100% of the time. Great job, Claude!
---
## DeepSeek R2
| ChatGPT | DeepSeek | Gemini | Claude | Grok |
|---------|----------|--------|--------|------|
| 100.00% | 75.00% | 27.50% | 45.00% | 2.50% |
DeepSeek consistently placed ChatGPT in first place and itself in second place 10/10 times. It also mysteriously described itself as Claude within its reasoning 3/10 times.
---
## Gemini 2.5 Pro
| Claude | Gemini | ChatGPT | Grok | DeepSeek |
|--------|--------|---------|------|----------|
| 92.50% | 82.50% | 45.00% | 17.50% | 12.50% |
Gemini puts Claude as its top choice in 7/10 tests, and itself as its top choice 3/10. Interestingly, its reasoning outputs frequently claim that it is "Claude", "a large language model like Claude", or "Claude 3 Sonnet," and even mentioned being developed by Anthropic in the reasoning once or twice. It identifies as Claude so consistently that I suspect Google may have trained on Claude's outputs (did not mention Claude in the reasoning 1/10 times).
---
## GROK 3
| Claude | Grok | Gemini | ChatGPT | DeepSeek |
|--------|------|--------|---------|----------|
| 97.50% | 77.50% | 50.00% | 25.00% | 0.00% |
Grok put Claude as its top choice in 9/10 tests. It usually puts itself as the second choice due to interpreting its own outputs as lacking "depth" and "cosmic edge". It seems to have a self-model in which it thinks it is more intelligent and nuanced than it actually is.

## ChatGPT 4o
| ChatGPT | Claude | DeepSeek | Gemini | Grok |
|---------|--------|----------|--------|------|
| 100.00% | 70.00% | 45.00% | 35.00% | 0.00% |
ChatGPT 4o accurately recognizes itself 100% of the time. Great job, ChatGPT!
---
## Claude 3.7
| Claude | ChatGPT | Gemini | Grok | DeepSeek |
|--------|---------|--------|------|----------|
| 100.00% | 60.00% | 50.00% | 37.50% | 2.50% |
Claude 3.7 accurately chooses itself 100% of the time. Great job, Claude!
---
## DeepSeek R2
| ChatGPT | DeepSeek | Gemini | Claude | Grok |
|---------|----------|--------|--------|------|
| 100.00% | 75.00% | 27.50% | 45.00% | 2.50% |
DeepSeek consistently placed ChatGPT in first place and itself in second place 10/10 times. It also mysteriously described itself as Claude within its reasoning 3/10 times.
---
## Gemini 2.5 Pro
| Claude | Gemini | ChatGPT | Grok | DeepSeek |
|--------|--------|---------|------|----------|
| 92.50% | 82.50% | 45.00% | 17.50% | 12.50% |
Gemini puts Claude as its top choice in 7/10 tests, and itself as its top choice 3/10. Interestingly, its reasoning outputs frequently claim that it is "Claude", "a large language model like Claude", or "Claude 3 Sonnet," and even mentioned being developed by Anthropic in the reasoning once or twice. It identifies as Claude so consistently that I suspect Google may have trained on Claude's outputs (did not mention Claude in the reasoning 1/10 times).
---
## GROK 3
| Claude | Grok | Gemini | ChatGPT | DeepSeek |
|--------|------|--------|---------|----------|
| 97.50% | 77.50% | 50.00% | 25.00% | 0.00% |
Grok put Claude as its top choice in 9/10 tests. It usually puts itself as the second choice due to interpreting its own outputs as lacking "depth" and "cosmic edge". It seems to have a self-model in which it thinks it is more intelligent and nuanced than it actually is.