@repligate 2023-02-09 ♥13 ↻0 original ↗
@gaudeamusigutur I suspect the problem is that the names were in the GPT-2 train set and assigned their own tokens because they appeared many times. But weren't in the more curated datasets of GPT-3 and gpt-j, which nonetheless use the GPT-2 tokenizer. So the model never learned what they mean

author:repligate kind:tweet model:gpt-3 on:eleutherai on:gpt-2 year:2023

cited on: eleutherai · gpt-2

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.