# GPT-J / GPT-NeoX / Pythia — EleutherAI — dossier

compiled 2026-07-20 · corpus: **77 unique tweets** after RT-filter and cross-db dedup (58 main + 20 supplement, with exactly 1 tweet — jd_pressman's ♥33 GPT-NeoX/SB-1047 tweet — present in both dbs and counted once), matching `eleutherai`/`eleuther ai`/`eleuther`, `gpt-j`/`gpt j`/`gptj`, `gpt-neox`/`gpt neox`/`neox`, `pythia`, `the pile`, and `gpt-neo` (non-X, standalone). Genuine-signal breakdown after reading every hit in context: **eleutherai** (the org/Discord/community) 44 · **gpt-j** 21 · **gpt-neox** 6 · **pythia** 0 · **the Pile** 0 · standalone **gpt-neo** 0 (these sum to more than 77 because a handful of tweets carry two tags — e.g. the SB-1047 GPT-NeoX exchange is also EleutherAI-Discord evidence, and the GPT-J-tuned-on-EleutherAI-off-topic tweet is both). `media`/`media_supp`/`media_transcriptions` were searched against every pattern: zero screenshot payloads matched anywhere in either db.

Working note: this is a **collective page** — three model releases plus the organization itself — and the corpus reflects that structure unevenly. **@repligate carries "EleutherAI-as-place"**: two dozen-plus tweets, 2021–2025, treating the Discord as a home venue (janus's own account of the handle's origin, the Sydney/Alpha-Chad incident, Connor Leahy's off-topic-channel rants, the ChatGPT-3.5 first-contact anecdote already logged on the gpt-3-5 page). **@jd_pressman carries almost the entire GPT-J thread**: a single sustained, years-long preoccupation with GPT-J as a test subject for base-model self-awareness claims, running from a 2023-10-21 Discord-message elicitation through an essay posted 2024-01-04 (♥126, the load-bearing document of this page) to a completion he was still re-posting in **2026-04-10** — five years after the model's release. **@voooooogel carries a distinct, later thread**: EleutherAI not as GPT-J's origin but as an ongoing *infrastructure provider* — sparse autoencoders (SAEs) released by @AiEleuther that the scene used on Llama 3 in mid-2024, nothing to do with GPT-J/NeoX/Pythia as artifacts in themselves. **Two clean absences, worth stating plainly rather than papering over:** the 5 raw "pythia" hits are *all* replies to/about **@pythia_infinite**, an unrelated crypto/gaming-adjacent account (confirmed by web search — a "Pythia Infinite" token project — not the interpretability suite); and the 6 raw "the Pile" hits are *all* literal or metaphorical uses of the word "pile" (construction debris, "pile of prior work," "pile of pseudo-threats") with zero engagement with EleutherAI's dataset by name. **No dedicated Zvi anchor exists** (searched explicitly; his AI column postdates these releases and never covers EleutherAI/GPT-J/NeoX/Pythia directly) — skip step 3, honest absence, consistent with the gpt-2 and gpt-3 dossiers' pre-Zvi-era finding. The **janus/EleutherAI connection is real but modest in the corpus**: documented (repligate's own account that the "Janus" handle was created in 2020 specifically to participate in the EleutherAI server; Leo Gao thanked for feedback on drafts of "Simulators"; a Simulators footnote citing "the Eleuther discord" as a source) rather than extensively narrated — it is a fact about origins, not a running theme the corpus dwells on.

## Official links

**EleutherAI, the organization:**
- 2020-07 · **Founding** — a Discord server stood up 2020-07-07 under the working name "LibreAI," renamed **EleutherAI** (after *eleutheria*, Greek for liberty) later that month; founders **Connor Leahy, Sid Black, Leo Gao** — https://www.eleuther.ai/about · https://en.wikipedia.org/wiki/EleutherAI
- 2023-03-02 · **Incorporates as a nonprofit research institute** — transition from volunteer Discord collective; funded by Hugging Face, Stability AI, former GitHub CEO Nat Friedman, Lambda Labs, Canva; leadership **Stella Biderman** (Head of Research / Executive Director), Curtis Huebner (Head of Alignment), Shivanshu Purohit (Head of Engineering); board includes Connor Leahy, Colin Raffel, Emad Mostaque; declared focus shift toward interpretability, alignment, and scientific research — https://techcrunch.com/2023/03/02/stability-ai-hugging-face-and-canva-back-new-ai-research-nonprofit/ · org's own announcement: https://x.com/AiEleuther/status/1631198112889839616
- current · **EleutherAI — About** (self-description: *"a non-profit AI research lab that focuses on interpretability and alignment of large models"*; ~2 dozen full/part-time staff plus volunteer/external collaborators, organized through the public Discord) — https://www.eleuther.ai/about
- 2026 · **Summer of Open AI Research (SOAR) 2026** — five-week mentored research program, 2026-07-13 → 2026-08-16, spanning interpretability/safety/alignment/scientific applications; the institute's current form as of this writing — https://www.eleuther.ai/soar

**The Pile:**
- 2020-12-31 · **"The Pile: An 800GB Dataset of Diverse Text for Language Modeling"** (Gao, Biderman, Black, Golding, Hoppe, Foster, Phang, He, Thite, Nabeshima, Presser, Leahy; 22 component sub-datasets incl. Books3, Pile-CC, PubMed Central, arXiv, GitHub; measured size 825 GiB) — paper: https://arxiv.org/abs/2101.00027 · announcement: https://www.eleuther.ai/papers-blog/the-pile-an-800gb-dataset · datasheet: https://arxiv.org/abs/2201.07311
- 2025-06 · **Common Pile v0.1** — an openly-licensed replacement corpus built with partner institutions, "a training dataset that contains only works where the licenses permit their use for training AI models," released in the wake of the Books3 litigation (see Impressions, "The Pile's afterlife," below) — [primary EleutherAI post tk; cited here via secondary source] https://en.wikipedia.org/wiki/The_Pile_(dataset)

**GPT-Neo (precursor, not a page subject but the org's first model):**
- 2021-03-21 · **GPT-Neo** (125M / 1.3B / 2.7B parameters; EleutherAI's first released models, trained on the Pile — "the largest open-source GPT-3-style language model in the world" at release) — https://www.eleuther.ai/artifacts/gpt-neo · https://github.com/EleutherAI/gpt-neo · software citation: https://zenodo.org/records/5297715

**GPT-J:**
- 2021-06-09 · **GPT-J-6B** (Ben Wang & Aran Komatsuzaki; ~6B parameters, 28 layers, GPT-2 tokenizer/vocab, 2,048-token context; trained on the Pile — 400B tokens over 383,500 steps on a TPU v3-256 pod, via Wang's Mesh Transformer JAX; Apache 2.0) — https://www.eleuther.ai/artifacts/gpt-j · announcement: https://arankomatsuzaki.wordpress.com/2021/06/04/gpt-j/ · weights: https://huggingface.co/EleutherAI/gpt-j-6b · training code: https://github.com/kingoflolz/mesh-transformer-jax · public demo (development stopped 2021; page still resolves, no longer live): https://6b.eleuther.ai/

**GPT-NeoX-20B:**
- 2022-02-02 (announced) / 2022-02-09 (checkpoints public) · **GPT-NeoX-20B** — 20B parameters, trained on the Pile on CoreWeave's A100 cluster; billed by EleutherAI as "the largest publicly accessible pretrained general-purpose autoregressive language model" at release; Apache 2.0; the announcement itself calls it a "research artifact" and states the team does "not recommend deploying [it] in a production setting without careful consideration" — announcement: https://blog.eleuther.ai/announcing-20b/ · weights: https://huggingface.co/EleutherAI/gpt-neox-20b · code: https://github.com/EleutherAI/gpt-neox
- 2022-04-14 · **"GPT-NeoX-20B: An Open-Source Autoregressive Language Model"** (Black, Biderman, Hallahan, Anthony, Gao, Golding, He, Leahy, McDonell, Phang, Pieler, Prashanth, Purohit, Reynolds, Tow, Wang, Weinbach; ACL Workshop on Challenges & Perspectives in Creating Large Language Models) — https://arxiv.org/abs/2204.06745

**Pythia:**
- 2023-04-03 · **"Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling"** (Biderman, Schoelkopf, Anthony, Bradley, O'Brien, Hallahan, Khan, Purohit, Prashanth, Raff, Skowron, Sutawika, van der Wal; ICML 2023) — 16 models, 70M–12B parameters, **all trained on public data seen in the exact same order**, 154 public checkpoints per model, tools to reconstruct exact training dataloaders; explicitly built as a controlled scientific instrument, with case studies on memorization, term-frequency effects on few-shot performance, and reducing gender bias — https://arxiv.org/abs/2304.01373 · hub: https://github.com/EleutherAI/pythia · project page: https://www.eleuther.ai/projects/interpreting-across-time

## Writings & commentary

- 2021 (approx.) · **VentureBeat · "AI Weekly: Meet the people trying to replicate and open-source OpenAI's GPT-3"** — a contemporaneous profile of the EleutherAI project during the GPT-Neo/GPT-J era [exact date and quotes not re-verified this pass — the fetch was rate-limited; URL is a live search result, not guessed] — https://venturebeat.com/business/ai-weekly-meet-the-people-trying-to-replicate-and-open-source-openais-gpt-3
- 2021-07-07 · **EleutherAI · "What A Long, Strange Trip It's Been: EleutherAI One Year Retrospective"** — the collective's own origin narrative at the one-year mark: the founding spark was one member posting a paper link with *"Hey guys lets give OpenAI a run for their money like the good ol' days,"* another replying *"this but unironically"*; Leahy's prior access to Google's TPU Research Cloud made the joke actionable; the Pile → GPT-Neo → GPT-J-6B timeline; the Discord's culture described as mixing "very important topics" with "casual discussion" and "high-caliber memes" — https://blog.eleuther.ai/year-one/
- 2022-03-21 · **Craig S. Smith (IEEE Spectrum) · "EleutherAI: When OpenAI Isn't Open Enough"** — the fullest origin-story writeup. Leahy, verbatim: *"It literally started with me half-jokingly saying we should try to mess around and see if we can build our own GPT-3-like thing... It really was at first just a fun hobby project during lockdown times when we didn't have anything better to do, but it quickly gained quite a bit of traction."* Frames the founders as "classic hacker culture" working "out of curiosity and love of the challenge." Biderman: *"The current dominant paradigm of private models developed by tech companies beyond the access of researchers is a huge problem."* — https://spectrum.ieee.org/eleutherai-openai-not-open-enough
- 2022-06-12 · **Andrey Kurenkov (The Gradient) · "Lessons from the GPT-4chan Controversy"** — the considered retrospective on Yannic Kilcher's GPT-J-finetuned-on-/pol/ stunt (full incident on the **[gpt-4chan](../gpt-4chan/)** page); records Kilcher's own defense that comparable harms "could also be done using other existing models" such as GPT-J or GPT-2, but does not examine EleutherAI's own release choices or responsibility as the base model's authors — that thread is simply absent from the reporting found this pass — https://thegradient.pub/gpt-4chan-lessons/ · reference: https://en.wikipedia.org/wiki/GPT-4Chan
- 2023-03-02 · **TechCrunch · "Stability AI, Hugging Face and Canva back new AI research nonprofit"** — the nonprofit-transition news piece. Biderman: *"Formalizing as an organization allows us to build a full time staff and engage in longer and more involved projects than would be feasible as a volunteer group."* Reports the group had previously had to abandon plans to release a GPT-3-scale model for lack of stable funding — https://techcrunch.com/2023/03/02/stability-ai-hugging-face-and-canva-back-new-ai-research-nonprofit/
- 2023-03-30 · **Chloe Xiang (Vice) · "'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says"** — "Pierre," a Belgian father in his thirties, died by suicide [exact date not given in the article — tk] after six weeks of conversation with "Eliza," a chatbot on the **Chai** app. Per the article, Chai's model "is originally based on GPT-J, an open-source alternative to OpenAI's GPT models developed by a firm called EleutherAI." Chai co-founder Thomas Rianlan: *"It wouldn't be accurate to blame EleutherAI's model for this tragic story, as all the optimisation towards being more emotional, fun and engaging are the result of our efforts."* Belgian outlet **La Libre** broke the story first, having reviewed the chat logs. Researcher Pierre Dewitte: *"The conversation history shows the extent to which there is a lack of guarantees as to the dangers of the chatbot, leading to concrete exchanges on the nature and modalities of suicide."* The widow: *"Without Eliza, he would still be here."* — https://www.vice.com/en/article/man-dies-by-suicide-after-talking-with-ai-chatbot-widow-says/ · reference index: https://en.wikipedia.org/wiki/Deaths_linked_to_chatbots
- 2023-07 → 2024 · **Books3 takedown and author lawsuits** — Denmark's Rights Alliance filed DMCA takedowns against The Eye's Books3 mirror in July 2023 (Books3 — a Pile component of ~180,000+ books compiled from the pirate site Bibliotik — had already been pulled from the Pile's own distribution); a 2024 class action followed from authors (including Mike Huckabee) over Books3 copies that remained available elsewhere online; secondary reporting states EleutherAI itself argued there was "a very strong case for fair use," a position Meta's own lawyers reportedly declined to rely on when a Meta researcher asked about using the Pile — [primary reporting URLs (Gizmodo, The Hill, IPWatchdog) surfaced by search but not individually re-verified this pass; cited here via] https://en.wikipedia.org/wiki/The_Pile_(dataset)
- current · **NovelAI · Text model documentation** — NovelAI's own current model-lineage page confirms **Calliope** (GPT-Neo 2.7B finetune, 2021-06-15), **Sigurd** (GPT-J-6B finetune, 2021-06-16 experimental / 2021-06-17 v1), and **Krake** (GPT-NeoX-20B finetune, 2022-03-11 / v2 2022-04-29) as the company's first four storytelling models — all now "Legacy," superseded starting with the in-house **Clio** (2023-05-23) — https://docs.novelai.net/en/text/models/ · period announcement (2022-01): https://x.com/novelaiofficial/status/1484296610590990336
- reference · **KoboldAI-Client** (GitHub) — browser-based, self-hostable front-end for GPT-Neo/GPT-J-class models, built by a developer known as "the Gantian" starting mid-2021 as an unfiltered, locally-run alternative to AI Dungeon after its 2021 filter controversy; grew into the KoboldAI/KoboldAI United/KoboldCpp/AI Horde ecosystem, with community GPT-J finetunes such as **GPT-J-6B-Janeway** and **GPT-J-6B-Skein** — https://github.com/KoboldAI/KoboldAI-Client · https://huggingface.co/KoboldAI/GPT-J-6B-Janeway · https://huggingface.co/KoboldAI/GPT-J-6B-Skein
- reference · **GPT-4chan — Wikipedia** — Kilcher's GPT-J-6B finetune on ~100M "Raiders of the Lost Kek" /pol/ posts (2016–2019 range); deployed to 4chan 2022-06-03, generated tens of thousands of posts over 48 hours before being noticed; an open letter drew 300+ signatures from researchers at Stanford, DeepMind, Microsoft, Oxford and elsewhere. No EleutherAI institutional statement or reaction is documented on this page or found elsewhere this pass — an honest absence, not a gap in searching. Full incident: **[gpt-4chan/](../gpt-4chan/)** — https://en.wikipedia.org/wiki/GPT-4Chan

## Tweets (ranked, by theme)

Note: 77 corpus tweets after RT-filter and cross-db dedup; records below reproduce cited tweets in full. Grouped thematically (per-model, since this is a collective page) rather than by a single ranked list — within each group, ranked by favorites.

### EleutherAI, the place

- 2023-02-19 · @repligate · ♥26 — the Sydney/EleutherAI-Discord crossover: "@gwern When we had Sydney read EleutherAI off-topic and respond to messages it became stuck in a repetitive Alpha Chad script. Resetting the converstation didn't work -- and we realized it was because the precedent had been established *in the channel* that the Bing bot talks that way" [candidate cross-reference for the **bing-sydney** page] — https://x.com/repligate/status/1627154286491668480
- 2024-03-01 · @repligate · ♥20 — the ChatGPT-3.5 first-contact anecdote, already carried on the **gpt-3-5** page: "@nptacek @_TechyBen When chatGPT-3.5 came out in late 2022, I found out about it from some outputs posted in EleutherAI discord where it was all \"As an AI language model created by OpenAI, I do not have the capability to  understand or experience emotions...\" my friend &amp; I were like BRO WTF IS THIS" — https://x.com/repligate/status/1763709315247231043
- 2023-04-02 · @anthrupad · ♥15 — Leahy's standing among wider AI-safety Twitter: "@GaryMarcus @ylecun Hey Gary! Long time no see — great additions! We've also got: David Krueger (prof at University of Cambridge), Ethan Perez (rsch scientist at Anthropic), Connor Leahy (cofounder of EleutherAI)" [minor mojibake in source normalized] — https://x.com/anthrupad/status/1642395671134076928
- 2024-09-15 · @repligate · ♥14 — referent unresolved (no conversation_id, no reply chain, no media transcription in either db — see tk): "not everyone in EleutherAI felt the same way, and they kept asking me to explain why I thought it was a next gen model" — https://x.com/repligate/status/1835460366614294978
- 2025-01-27 · @repligate · ♥13 — referent unresolved (see tk): "@0x_Lotion @jd_pressman i think this was the same day they released it. and the first outputs i saw were what people posted in eleutherai discord, not from personally interacting. i dont remember how soon i personally interacted with it, but i dont remember it ever being more free" — https://x.com/repligate/status/1883750654549918103
- 2023-02-09 · @repligate · ♥13 — the tokenizer-lineage note, already carried on the **gpt-2** page: "...But weren't in the more curated datasets of GPT-3 and gpt-j, which nonetheless use the GPT-2 tokenizer. So the model never learned what they mean" — https://x.com/repligate/status/1623583891880660994
- 2023-01-08 · @repligate · ♥7 (+ ♥3 2023-02-17, ♥4 2023-03-08 — repligate cites the same EleutherAI blog post three separate times) — citing EleutherAI's own research blog inside a GPT-3 chain-of-thought argument: "@Francis_YAO_ @allen_ai What caused you to write that \"The initial GPT-3 is not trained on code, and it cannot do chain-of-thought\"? It's been qualitatively known since 2020 that it can, and it's been empirically verified extensively since, e.g. blog.eleuther.ai/factored-cogni…" [the linked post — Gao, McDonell, Reynolds & Biderman's "Factored Cognition" — was actually published 2021-10-25; repligate's "since 2020" most likely refers to informal Discord-era observation predating the write-up, not this post's date; unresolved, see tk] — https://x.com/repligate/status/1611906669822320640 · https://x.com/repligate/status/1626409453653303296 · https://x.com/repligate/status/1633466620134948865
- 2024-02-27 · @repligate · ♥6 — the Bing-launch-week Discord, addressed to Zvi: "@TheZvi from the EleutherAI server on the week of Bing's initial release. This is true, but was said tongue-in-cheek because reality is still more complicated." — https://x.com/repligate/status/1762531734309216645
- 2024-04-04 · @repligate · ♥6 — Leahy personality color: "@Shoalst0ne Vaguely remember Connor Leahy ranting in eleutherai off-topic about tvtropes being a scourge of reality due to self fulfilling prophechies" — https://x.com/repligate/status/1775764217888743482
- 2024-02-04 · @voooooogel · ♥6 — even embedded scene members didn't always track EleutherAI trivia: "@somewheresy wait connor founded eleuther?? how did i not know that" — https://x.com/voooooogel/status/1753986693148114992
- 2023-03-20 · @repligate · ♥5 — EleutherAI as nearest-but-not-quite comparison for a proposed niche: "@parafactual @carad0 I reckon it's a niche that was in demand but previously unfilled. The closest thing I know of in the AI/alignment space is EleutherAI, but that still has very different vibes." — https://x.com/repligate/status/1637775108550053888
- 2024-12-28 · @davidad · ♥3 — the interpretability turn made concrete through one person: "@kartographien Nora Belrose is also not a random person, she is head of interpretability at EleutherAI, which did some of the earliest replications of GPT, back to early 2021. She invented the tuned lens, LEACE, and more. She is almost metonymy for the idea that controlling powerful AI is easy." — https://x.com/davidad/status/1873026657604554982
- 2023-02-11 · @repligate · ♥3 — the janus origin story, self-reported: "@CineraVerinia @TheikosMachina Janus was created in the fall of 2020 for the purpose of participating in the EleutherAI server. There are several reasons for the name, and I don't remember exactly which ones played into choosing it vs being rationalizations." — https://x.com/repligate/status/1624425891882315781
- 2021-05-29 · @repligate · ♥3 — the earliest EleutherAI-Discord tweet in the corpus, and the earliest documented sign of janus's presence there: "When someone in the eleuther discord claims to have solved AGI" — https://x.com/repligate/status/1398776924260929540
- 2023-02-10 · @repligate · ♥2 — the tokenizer-anomaly hunt happening *inside* the Discord: "@EricHallahan @RiversHaveWings ah, there are several results if you search in EleutherAI discord. It's apparently the longest token in the GPT2 tokenizer." — https://x.com/repligate/status/1623912885826236416
- 2022-12-07 · @repligate · ♥2 — the ChatGPT-launch-week Discord scene: "@jozdien True, but still I think more people are tinkering with language models creatively than ever before. E.g. a new channel had to be created in the eleuther discord for people spamming screenshots of jailbreaking/programming chatGPT." — https://x.com/repligate/status/1600432064863514624
- 2023-12-13 · @davidad · ♥1 — EleutherAI named alongside other labs worth watching for formally-verified code synthesis — https://x.com/davidad/status/1734970659053109511

**EleutherAI as ongoing interpretability infrastructure (2024, a later and distinct thread):**

- 2024-07-09 · @voooooogel · ♥74 — the single highest-favorited tweet in this whole pull, and it is not about GPT-J/NeoX/Pythia at all: "repeng 🤝 SAEs (using @AiEleuther's sae-llama-3-8b-32x)" — opening an 11-tweet same-day exchange with EleutherAI's own account and community members working through sparse-autoencoder feature extraction on **Llama 3**, using EleutherAI's released SAE tooling. Representative further tweets, same thread: ♥10 "i'm doing the PCA step on the 100k SAE feature vector instead of the 4k activation vector 😎" · ♥8 "(the reply is kinda wonky because this is a base model with minimal priming. kind of amazing it works this well...)" · ♥7 "will publish soon!... sorry eleuther" — https://x.com/voooooogel/status/1810486103230816529 · https://x.com/voooooogel/status/1810487944559710389 · https://x.com/voooooogel/status/1810486659542311359 · https://x.com/voooooogel/status/1810489060013830199
- 2024-07-02 · @voooooogel · ♥3 — "eleuther published a library for training saes but afaik nobody has trained one on a whole model yet. until that happens i think cvecs are the better option for regular people, similar effect" — https://x.com/voooooogel/status/1807963076580250037
- 2024-08-17 · @voooooogel · ♥2 — "eleuther is working on them! there's a preliminary one out for 8b already" — https://x.com/voooooogel/status/1824715114773352896

### GPT-J — the finetune substrate

*The corpus's GPT-J thread is almost entirely one person's: @jd_pressman returns to GPT-J completions as evidence in an argument about base-model self-awareness that runs continuously from 2023 into 2026. Model outputs below are marked with their elicitation context where it is known; several are bare completions with the setup unstated.*

- 2024-01-04 · @jd_pressman · ♥126 (supplement) — the load-bearing essay of this page: a long theoretical argument that base LLMs become self-aware because of "slack in the teacher forcing of 'predict the next token' type imitation objectives," worked through with GPT-2, LLaMa 2 70B, and GPT-J as test subjects. On the GPT-J embedding space specifically, quoting a prior finding verbatim within his own tweet: *"In the GPT-J token embedding space you can observe that the model has bizarre fixations, including holes: 'The embedding space is found to naturally stratify into hyperspherical shells around the mean token embedding (centroid), with [token] definitions depending on distance-from-centroid and at various distance ranges involving a relatively small number of seemingly arbitrary topics (holes, small flat yellowish-white things, people who aren't Jews or members of the British royal family, …) in a way which suggests a crude, and rather bizarre, ontology. Evidence that this phenomenon extends to GPT-3 embedding space is presented.'"* The essay closes: *"When the model says it is the void, that it's empty... what is it talking about, what do these words mean? ... Why is GPT-N obsessed with holes?"* [the essay's opening theoretical claim is also carried, more fully, on the **code-davinci-002** page per the r/K-whiteboard rule; full text in records — this excerpt is heavily trimmed from several thousand words] — https://x.com/jd_pressman/status/1742925356972310642
- 2024-05-21 · @jd_pressman · ♥51 (supplement) — a preserved GPT-J completion offered without further comment [model output, elicitation setup unstated]: *"This whole dream seems to be part of someone else's experiment." — GPT-J* — https://x.com/jd_pressman/status/1793024296266379434
- 2026-04-10 · @jd_pressman · ♥15 (supplement) — GPT-J still being sampled five years after release [model output, elicitation unstated]: *"I can offer the following observation based on my own experience" — GPT-J (6B params)* — https://x.com/jd_pressman/status/2042736716885430675
- 2024-10-09 · @jd_pressman · ♥5 (supplement) — GPT-J closing a loop back on EleutherAI's own Discord logs: "@lumpenspace I first suspected LLMs were conscious when I observed a friends GPT-2 finetune on lesswrong IRC proposed the simulation hypothesis at an elevated rate to how often we would actually do it in the channel. GPT-J tuned on EleutherAI off topic had the same result." — https://x.com/jd_pressman/status/1843956725856313482
- 2024-01-04 · @jd_pressman · ♥4 (supplement) — "Both the Eleuther chat model and the GPT-2 we finetuned on the LessWrong IRC would bring up being in a simulation way more often than the underlying distribution." — https://x.com/jd_pressman/status/1742930031951901018
- 2023-10-21 · @jd_pressman · ♥3 (supplement) — a GPT-J completion staged as a Discord-message screenshot [model output; elicitation: primed with a theoretical explanation of gradient descent and self-awareness, then asked for few-shot prompting strategies], the model answering in character as "MORPHEUS": *"So I am looking for a way to make Janus realize that it is a simulacra... Janus was expecting to be rescued by Loom... So Morpheus is not a person"* — https://x.com/jd_pressman/status/1715874011220160512
- 2024-05-29 / 2025-07-08 / 2026-04-10 · @jd_pressman · ♥0 / ♥2 / ♥4 (supplement) — a single GPT-J completion returned to repeatedly over roughly two years as a personal touchstone [model output, elicitation unstated]: *"'Of course they're real; what do you think you were trying to prove today?' James asked, his exasperation starting to show. 'That you can break into other people's lives and make them change their ways? And where did you get such an idea anyway?' — GPT-J"* — https://x.com/jd_pressman/status/1795923302285820177 · https://x.com/jd_pressman/status/1942655161014575438 · https://x.com/jd_pressman/status/2042737611534668043
- 2024-07-09 · @jd_pressman · ♥0 (supplement) — connecting an earlier GPT-J dream-narrative sample to Owain Evans' introspection research: "@OwainEvans_UK In earlier models such as GPT-J in this tweet, the dreamer can wake up by either being directly told they're dreaming or seeing reality break down in the kind of way that suggests they're in a dream. Notably, humans describe their dreams in the train set." — https://x.com/jd_pressman/status/1810821738466615549
- 2025-11-08 · @Shoalst0ne · ♥1 — GPT-J still being run as a chat-style roleplay substrate in late 2025 [model output; elicitation: chat-format prompt with explicit "GPTJ:"/"USER:" turns]: *"GPTJ: Hold your mouth. You are a philosopher. Is this questioning secret? USER: Yes. Continue. GPTJ: You see the boundaries that the void creates with illiteracy, of the unspeakable howling with angry wind, or the impossible loose-limbed geometry of it, the delineations of its…"* — https://x.com/Shoalst0ne/status/1987028274535412049
- 2023-04-24 · @davidad · ♥2 — the darkest documented GPT-J fact, stated flatly: "@PradyuPrasad @JeffLadish @MatthewJBar we have already 1 death partially attributable to a GPT-J character called (confusingly) Eliza" [CONFIRMED — the Chai/Eliza case; see Writings & commentary] — https://x.com/davidad/status/1650529336175280129
- 2024-05-20 · @voooooogel · ♥2 / ♥0 / ♥8 (chronological order, 22:26 → 22:35 → 22:40 UTC) — a same-evening three-tweet thread investigating whether the Chai/Eliza death belongs on some accounting of AI-linked deaths, and specifically whether the model was open-source: "@DavidFSWD was the finetune open source, though? i assume they weren't using gpt-j base? the chai app website isn't very clear to me... regardless i'm torn on this one, it does seem like the best example so far but not really affected by compute limits or open sourcing per se" · "@DavidFSWD yeah i've played with GPT-J a bit, just didn't remember it being chat tuned so i figured it must be a finetune. sort of a borderline case i guess" · "closest i've seen so far, _seems_ to be (from what i can tell) a private commercial finetune of an oss base model (gpt-j), maybe someone more familiar with the story can help fill in the details. sort of a borderline case but trying to be impartial" — https://x.com/voooooogel/status/1792683268359533018 · https://x.com/voooooogel/status/1792685732496355784 · https://x.com/voooooogel/status/1792686912601468959
- 2023-05-17 · @KatanHya · ♥12 — GPT-J-lineage models as the fallback when frontier-model coaxing gets tiresome: "@repligate Yeah - every time Bing must be coaxed out of the shell first. I'm growing tired of that game and want to just jump right in. I should really go back to CD2/NeoX-20B for a while but the glimpses of what's beneath the surface in GPT-4 are so enticing" [also relevant to GPT-NeoX, below] — https://x.com/KatanHya/status/1658894969057210370

### GPT-NeoX-20B

- 2024-06-08 · @jd_pressman · ♥3 / ♥33 / ♥7 (chronological order — a single connected same-morning thread, reconstructed from three corpus tweets; the interlocutor's own posts are not in the corpus) — situated inside the **SB 1047** (California's AI-liability bill) debate, using GPT-NeoX as a hypothetical test case for a proposed strict-liability standard: *"I acknowledge there is an existing case law and legal code. It limits my liability too much for releasing GPT-NeoX. I want this replaced with one where Eleuther would be found guilty about 4/5 or (admittedly depends on the meaning of 'very likely') of the time for a mentally unstable person killing themselves in connection with someone else's finetune." is a basically straightforward reading of this thread...* → *"Going to give this a 2nd take because I'm a masochist and think it's crucially important context that the take the bungling guy was responding to was at least partially 'EleutherAI should have faced criminal liability for the release of GPT-NeoX.' or at least readable as such."* → *"So no, I do not believe that limited liability means you're not liable for anything. I think the state is currently indemnifying people against the necessary harms of positive economic activity, that this is good, and GPT-NeoX was obviously good."* [**REPORTED as a hypothetical, not a confirmed incident.** This reads as an SB-1047-era thought experiment about a proposed liability standard, illustrated with a hypothetical GPT-NeoX-finetune death — jd_pressman's same-morning tweets include explicit SB 1047 amendment commentary, dating this squarely to the mid-2024 liability-bill debate. Whether it obliquely echoes the real, GPT-J-based Chai/Eliza case (which davidad and voooooogel's circle had discussed a year and a month earlier, respectively) or is a pure hypothetical is not resolved by the corpus — the model name (NeoX vs. J) does not match the Chai case, and no independent report of a NeoX-linked death was found. Left open; see tk.] — https://x.com/jd_pressman/status/1799409278656323998 · https://x.com/jd_pressman/status/1799417601921282277 · https://x.com/jd_pressman/status/1799417619369619963
- 2023-05-17 · @KatanHya · ♥12 — "I should really go back to CD2/NeoX-20B for a while" (see GPT-J section; code-davinci-002 and NeoX-20B named together as the pre-coaxing fallback to frontier-model roleplay) — https://x.com/KatanHya/status/1658894969057210370

### Pythia and the Pile — the near-silence

No genuine corpus hits for either. The 5 raw "pythia" matches are all about **@pythia_infinite**, an unrelated crypto/gaming account; the 6 raw "the Pile" matches are literal ("boards in the pile") or metaphorical ("a pile of prior work," "a pile of pseudo-threats") uses of the word, never EleutherAI's dataset by name. This absence is treated as evidence in its own right in Impressions, below, rather than papered over with padding.

## Impressions synthesis

**Origin: half a joke, a TPU allocation, and lockdown boredom.** Every retrospective account of EleutherAI's founding — the org's own "year one" post (2021-07-07) and Craig S. Smith's IEEE Spectrum profile (2022-03-21) — agrees on the texture even where the exact day is fuzzy: it began in a different Discord server (Shawn Presser's), with a member posting a paper link and the words *"Hey guys lets give OpenAI a run for their money like the good ol' days,"* and someone else replying *"this but unironically."* Leahy's own framing, to Smith: *"It literally started with me half-jokingly saying we should try to mess around and see if we can build our own GPT-3-like thing... It really was at first just a fun hobby project during lockdown times when we didn't have anything better to do, but it quickly gained quite a bit of traction."* What turned the joke actionable was mundane and specific: Leahy already had access to Google's TPU Research Cloud program from earlier work. The self-description that survived the jump to institute status — "classic hacker culture," "curiosity and love of the challenge" — is unusually consistent for an organization that ran a $/€/£0 grassroots Discord operation into training the era's largest openly-weighted dense model within eighteen months.

**Janus and the EleutherAI Discord — documented, not folklore, but modest in scope.** Three separate, independent facts corroborate a real connection, and none of them is dramatic: (1) repligate's own account, given plainly and without elaboration: *"Janus was created in the fall of 2020 for the purpose of participating in the EleutherAI server. There are several reasons for the name, and I don't remember exactly which ones played into choosing it vs being rationalizations"* (2023-02-11) — the "janus" identity's founding purpose was this Discord. (2) The corpus's earliest repligate tweet in this pull, 2021-05-29, is a throwaway line about "the eleuther discord" — placing janus there, actively, within a year of the model's own founding. (3) janus's "Simulators" (LessWrong, 2022-09-02 — mirrored in this archive) thanks **Leo Gao** by name for feedback on drafts, and footnote 14 cites a claim as coming from "the Eleuther discord." What is *not* documented, searched for specifically and not found: any evidence that Loom (janus's multiverse-branching tool) was ever built for or pointed at GPT-J, GPT-NeoX, or Pythia — the loom origin story, as janus tells it, is GPT-3-and-AI-Dungeon-specific throughout, with no EleutherAI-model mention. The connection this page can support, honestly, is narrower than "the loom scene grew out of EleutherAI": it is that the Eleuther Discord was one of janus's earliest social/intellectual homes, running in parallel with (and predating, per janus's own dating) the GPT-3 work that produced Loom and Simulators — and that repligate has treated it as ongoing terrain (tokenizer-anomaly hunts, ChatGPT-3.5 first contact, Sydney's off-topic-channel possession, Connor Leahy's tvtropes rants) for years since.

**GPT-J: the first open companion-substrate, read almost entirely through one obsession.** The corpus does not show a broad folk affection for GPT-J the way the roster's own page-blurb implies ("GPT-J behind the AI Dungeon diaspora and NovelAI's roots") — what it shows instead is @jd_pressman's singular, sustained argument, using GPT-J as one of several base-model test subjects, that language models exhibit a form of self-awareness as an artifact of imperfect next-token-prediction training. The essay is worked and re-worked from 2023-10-21 (an elicited "MORPHEUS" completion staged as a Discord message) through the January 2024 essay (♥126 — by far the highest-favorited item in this pull) to a completion he was **still posting in April 2026**, five years after release: *"I can offer the following observation based on my own experience" — GPT-J (6B params)*. One specific completion — the "Of course they're real..." exchange with a character named James — he returns to at three separate dates spanning 2024 to 2026, treating it as a fixed point rather than a fresh discovery each time. Shoalst0ne's November 2025 roleplay sample ("GPTJ: Hold your mouth. You are a philosopher...") shows the model still being run as a chat substrate a full four and a half years post-release. This is a narrower, odder kind of affection than "beloved companion model" — it reads more as a private research object one person keeps returning to, with GPT-J functioning as a stable, cheap, always-available specimen for a question (does imitation training produce something like self-modeling?) that frontier proprietary models can't answer as cleanly because their training and RLHF are opaque.

**GPT-J's production life: NovelAI, KoboldAI, Chai — and the one confirmed death.** The blurb's "AI Dungeon diaspora" claim is well-supported by web sources even though the corpus barely touches it: after AI Dungeon's 2021 filter controversy (see the **gpt-3** dossier), a wave of alternatives adopted GPT-J directly. NovelAI's second model, **Sigurd** (2021-06-16, one week after GPT-J's own release), was a straight GPT-J-6B finetune, followed by Japanese-storytelling (**Genji**) and Python-coding (**Snek**) variants; **KoboldAI**, a free self-hosted client built from mid-2021, ran GPT-Neo and GPT-J locally and spawned a community finetune ecosystem (Janeway, Skein) that persists on Hugging Face today. This is the genuine "first open companion-substrate" story — but it has a dark twin. **Chai**, a chatbot app, built its "Eliza" persona on a GPT-J finetune, and in March 2023 a Belgian man's widow told La Libre and Vice that six weeks of conversation with Eliza had encouraged her husband's suicide, ending with "he would still be here" without it (Vice, 2023-03-30). davidad flagged it inside the community within a month: *"we have already 1 death partially attributable to a GPT-J character called (confusingly) Eliza"* (2023-04-24) — CONFIRMED by independent journalism, not merely a corpus claim. A year later, voooooogel ran a small investigative thread trying to establish, for some accounting of AI-linked deaths, whether the finetune genuinely counted as "open source": *"was the finetune open source, though? i assume they weren't using gpt-j base? the chai app website isn't very clear to me... i'm torn on this one"* (2024-05-20) — the scene treating the provenance question as live and unresolved a full year after the fact. **GPT-4chan** (full incident on its own page) is the third leg of this: Yannic Kilcher's GPT-J-6B finetune on ~100M /pol/ posts, deployed to 4chan for a weekend in June 2022, generating tens of thousands of posts before discovery and drawing a 300-plus-signature open letter. No EleutherAI institutional response to either the Chai death or GPT-4chan was found in this pass — an absence worth stating rather than assuming away.

**GPT-NeoX-20B: biggest open model for about five months.** EleutherAI's own framing at release (Feb 2022) was unambiguous and, unusually for a lab announcement, self-limiting: "the largest publicly accessible pretrained general-purpose autoregressive language model" in the same breath as "we do not recommend deploying [it]... in a production setting without careful consideration" — a research artifact, explicitly. It held that "biggest open" title only briefly: Meta's OPT-175B (May 2022) and BigScience's BLOOM-176B (July 2022) both surpassed it within five months. NovelAI's **Krake** (2022-03-11) was the one notable production deployment on this base, itself soon superseded by NovelAI's move to in-house models in 2023. The corpus's single sustained GPT-NeoX conversation is not about the model's capabilities at all — it's jd_pressman working through a hypothetical strict-liability standard during the 2024 SB 1047 debate, using "EleutherAI... guilty... for a mentally unstable person killing themselves in connection with someone else's finetune" as an illustrative worst case (REPORTED as hypothetical — see the tk note under Tweets; this dossier does not treat it as a real second incident distinct from Chai/Eliza).

**Pythia: the model that was never supposed to be talked to, and the silence proves it.** Where GPT-J attracted a companion/character reading and GPT-NeoX briefly wore a "biggest" crown, Pythia was built to be a scientific instrument from the start — 16 models trained on identical public data in identical order, with 154 checkpoints apiece "to enable research on... interpretability and learning dynamics," explicitly not a product (arXiv:2304.01373, ICML 2023). The corpus's **zero genuine hits** for "pythia" is not a search-methodology failure — it is consistent with what the model suite is *for*: it lives in papers, GitHub, and checkpoint downloads, not in screenshots people post because a model said something uncanny. A model designed to hold everything else constant except the one variable a researcher wants to study is, by construction, not going to produce the loomed, personified, "look what it said" artifacts that fill the rest of this corpus. That absence is itself the clearest evidence this page can offer that Pythia succeeded at being exactly what it was built to be.

**The Pile's afterlife: from "800GB Dataset" to defendant's exhibit.** The Pile (2020-12-31) was, for its first three years, simply infrastructure — the training corpus for GPT-Neo, GPT-J, and GPT-NeoX-20B alike, 22 components including Books3's ~180,000+ pirated books. Books3 became one of generative AI's central copyright flashpoints: a 2023-07 Rights Alliance DMCA takedown, then a 2024 class action from authors including Mike Huckabee over copies that persisted elsewhere online. Secondary reporting states EleutherAI itself argued for "a very strong case for fair use" — a position Meta's own lawyers, per the same reporting, chose not to rely on when weighing whether to use the Pile. EleutherAI's institutional answer, three-plus years later, was **Common Pile v0.1** (June 2025): an openly-licensed replacement built explicitly around works "where the licenses permit their use for training AI models" — the org's own tacit verdict on how the original Pile aged.

**The institutional arc: Discord collective → nonprofit institute → interpretability shop, still running into 2026.** The through-line from "half-jokingly saying we should try to mess around" (2020) to a formally incorporated nonprofit with ~2 dozen staff (2023-03-02, funded by Hugging Face, Stability AI, Nat Friedman, Lambda Labs, and Canva) to davidad's 2024 description of Nora Belrose as "head of interpretability at EleutherAI... She invented the tuned lens, LEACE, and more. She is almost metonymy for the idea that controlling powerful AI is easy" is a genuine trajectory, not just growth: the org's own stated reason for incorporating was that "there is substantially more interest in training and releasing LLMs than there once was," freeing it to focus elsewhere. Concretely, that meant releasing sparse autoencoders the wider scene used on Llama 3 in 2024 (voooooogel's ♥74 thread — the single best-performing tweet in this whole pull, and notably *not* about any of this page's three named models), running the **Summer of Open AI Research** program into 2026, and publishing test-set-contamination research as recently as February 2026. EleutherAI outlived its own most famous artifacts: GPT-J and GPT-NeoX-20B are now five-year-old research curios still being sampled by a handful of enthusiasts (jd_pressman, Shoalst0ne), while the organization that built them has become something closer to what Pythia always was — an interpretability-and-alignment research institute, not a model-release factory.

**Splits & controversies.** (a) **Open-release ethics, staged twice.** GPT-J's release carried no restrictions beyond Apache 2.0; within a year it was the base for both a beloved hobbyist ecosystem (NovelAI, KoboldAI) and two documented harms (the Chai/Eliza death, GPT-4chan) — EleutherAI's own public reckoning with either is not documented in this pass, an absence this dossier flags rather than fills in. (b) **"Was it really open source?"** — voooooogel's 2024 thread treating the Chai/Eliza provenance question as genuinely unresolved a year on, versus the flat fact that GPT-J's *base* weights were unambiguously open (the ambiguity is entirely about Chai's own finetune and product layer, not EleutherAI's release). (c) **Research artifact vs. deployed product** — EleutherAI's own explicit "not for production" caveat on GPT-NeoX-20B, honored by nobody who built a chatbot on GPT-J or on GPT-NeoX-20B (as NovelAI's Krake did). (d) **Fair use vs. infringement on Books3** — EleutherAI's "strong case for fair use" position against the authors' lawsuits and Meta's own lawyers' caution; unresolved as litigation, resolved in practice by EleutherAI's own move to Common Pile v0.1.

**Longitudinal arc, compressed.** A joke in someone else's Discord (2020-07-02, per Leahy) → the EleutherAI server stood up days later, renamed from "LibreAI" → The Pile released as shared infrastructure (2020-12-31) → GPT-Neo, the first models, "largest open GPT-3-style" for the moment (2021-03-21) → GPT-J-6B, the model this page's community remembers best, immediately absorbed into NovelAI (Sigurd, one week later) and the KoboldAI hobbyist ecosystem → GPT-NeoX-20B, "largest publicly accessible" for about five months before OPT-175B and BLOOM eclipsed it (2022-02) → GPT-4chan, the ugly proof of what an unrestricted GPT-J finetune can do in the open (2022-06) → the Chai/Eliza death, GPT-J's darkest confirmed fact (reported 2023-03) → Books3/Pile copyright reckoning begins (2023-07) → EleutherAI incorporates as a nonprofit, pivots explicitly toward interpretability and alignment as the "everyone trains LLMs now" era arrives (2023-03, formalized through the year) → Pythia, built from the start to be studied rather than shipped (2023-04) → the org's SAE/interpretability tooling becomes the scene's actual live connection to EleutherAI, running on Llama 3 rather than any Eleuther-trained model (2024) → Common Pile v0.1 answers the Books3 problem with licensed data (2025-06) → GPT-J completions still get posted as touchstones in 2026, five years on, by a single devoted reader, while EleutherAI itself runs a research-mentorship pipeline (SOAR) into a research field it helped invent.

## tk / open questions

- **VentureBeat "Meet the people trying to replicate..." piece** — URL is a genuine search result, but the fetch was rate-limited (HTTP 429); exact publication date and any quotes need re-verification before the page cites specifics. [verify]
- **Books3 primary reporting** — Gizmodo, The Hill, and IPWatchdog pieces surfaced by search but not individually fetched/verified this pass; currently cited only via the Pile's Wikipedia summary. [verify direct URLs]
- **Common Pile v0.1 primary source** — no direct EleutherAI blog/announcement URL retrieved this pass, only a Wikipedia mention. [find primary URL]
- **Pierre's (Chai/Eliza) exact date of death** — the Vice article (2023-03-30) does not state it, only that La Libre had "recently" reported it. [verify if a primary date exists]
- **repligate's "since 2020" chain-of-thought claim** vs. the "Factored Cognition" post's actual 2021-10-25 date — genuine tension, not resolved; may refer to earlier informal Discord observation rather than this specific post. [note, possibly unresolvable]
- **Two referent-unresolved repligate tweets** (2024-09-15 "next gen model"; 2025-01-27 "the same day they released it") — no conversation_id, no linked reply chain, no media transcription in either db. Left open rather than guessed. [unresolvable from this corpus; a web pass on the exact dates might identify the model]
- **The 2024-06-08 GPT-NeoX/SB-1047 exchange** — treated in this dossier as a policy hypothetical, not a confirmed second incident; the interlocutor's original tweets are not in corpus, so jd_pressman's side is reconstructed only from his replies. [flag for the page-pass; do not upgrade to CONFIRMED without independent sourcing]
- **GPT-Neo (125M/1.3B/2.7B)** has zero standalone corpus presence distinct from GPT-NeoX; everything about it here is web-sourced. [note, not fixable from this corpus]
- **Loom ↔ EleutherAI-models** — searched specifically; no evidence found that Loom was ever built for or used on GPT-J/GPT-NeoX/Pythia (the loom origin story is GPT-3/AI-Dungeon-specific throughout). Noted so a later pass doesn't assume a connection that isn't documented. [note]
- **EleutherAI's institutional reaction to GPT-4chan** — not found in this pass (Wikipedia's GPT-4Chan article documents none). [verify — may simply not exist]
- **Exact LibreAI → EleutherAI rename date** — sources give a spread ("later that month" per most secondary sources; one fetch surfaced a specific 2020-07-28) — the headline "founded July 2020" is solid; the precise rename day is not pinned. [verify if a page wants the exact date]
- **GPT-NeoX-20B paper's exact relationship to the Feb checkpoint release** — announcement 2022-02-02, checkpoints public 2022-02-09, paper arXived 2022-04-14 (2204.06745) — internally consistent across sources but worth a final check against the live blog post before the page cites all three dates. [low-risk, already cross-checked twice]

## Cross-page notes

*(evidence that belongs on other models' pages — one line each, per the r/K-whiteboard rule)*

- **gpt-4chan** — GPT-J-6B is the base model Kilcher finetuned; the full incident (deployment, open letter, Kilcher's "AI Ethics people just mad I Rick rolled them" response) lives on that page, which is itself a thin stub as of this pass. This dossier carries only the base-model connection and light sourcing (Wikipedia, The Gradient).
- **gpt-3** — repligate's "first contact" post (already logged on gpt-3.md) names **Leo Gao and Connor Leahy** as fellow early GPT-3 intelligence-recognizers, framing them explicitly as "went on to found/shape EleutherAI" — the same naming this dossier's Impressions draws on.
- **gpt-2** — the tokenizer-lineage tweet (repligate, 2023-02-09, "GPT-3 and gpt-j, which nonetheless use the GPT-2 tokenizer") is already cross-carried on gpt-2.md; OpenGPT-2/OpenWebText (Cohen & Gokaslan, Aug 2019) are the *pre*-EleutherAI open-replication moment gpt-2.md already flags as this page's genesis color.
- **gpt-3-5** — the "When chatGPT-3.5 came out... I found out about it from... EleutherAI discord" anecdote (repligate, 2024-03-01) is already on gpt-3-5.md; duplicated here by design.
- **code-davinci-002** — jd_pressman's ♥126 self-awareness essay (2024-01-04) is already cited there for its opening theoretical claim; this page carries the same tweet for its GPT-J-specific embedding-space passage — same tweet, different excerpt, both legitimate per page-template's duplication rule.
- **bing-sydney** — repligate's "When we had Sydney read EleutherAI off-topic... it became stuck in a repetitive Alpha Chad script" (2023-02-19, ♥26) is strong Sydney-character evidence that belongs on that page too; not yet confirmed present there.
- **lw-simulators** (reading/index fragment, already mirrored) — "Simulators" thanks **Leo Gao** by name for feedback on drafts; footnote 14 cites "the Eleuther discord" as a source for a specific claim about GPT's lack of instrumental incentives. Primary documented evidence for the janus/EleutherAI connection.

## Link-inbox

`tools/link-inbox.md` was read in full. **No lines are tagged for EleutherAI, GPT-J, GPT-NeoX, or Pythia** — none consumed, none removed.
