GPT-J / GPT-NeoX / Pythia
EleutherAI began in July 2020 as a Discord collective formed to build open replications of GPT-3. It released The Pile (December 2020), GPT-Neo (March 2021), GPT-J-6B (June 2021), GPT-NeoX-20B (February 2022), and the Pythia suite (April 2023); GPT-J became the base model for NovelAI, KoboldAI, Yannic Kilcher’s GPT-4chan, and the Chai ‘Eliza’ chatbot linked to a 2023 suicide. The group incorporated as a nonprofit research institute in March 2023 and shifted toward interpretability and alignment; The Pile’s Books3 component became a copyright flashpoint, answered in 2025 by the licensed Common Pile v0.1.
Sources
Curated. Full compilation: dossier (77 corpus tweets after RT-filter and cross-db dedup). This is a collective page — three model releases plus the organization — and the corpus covers it unevenly: @repligate carries EleutherAI-as-Discord, @jd_pressman almost the entire GPT-J thread, @voooooogel a later interpretability-infrastructure thread; Pythia and The Pile have no genuine corpus presence at all (see Impressions). A known lens, not a neutral sample.
Official
- 2020-07 EleutherAI — About — founding: a Discord server stood up 2020-07-07 (working name “LibreAI,” renamed EleutherAI, after eleutheria, Greek for liberty); founders Connor Leahy, Sid Black, Leo Gao. Self-describes as “a non-profit AI research lab that focuses on interpretability and alignment of large models.”
- 2020-12-31 The Pile: An 800GB Dataset of Diverse Text for Language Modeling — 22 component sub-datasets (Books3, Pile-CC, PubMed Central, arXiv, GitHub); measured size 825 GiB. announcement.
- 2021-03-21 GPT-Neo — 125M / 1.3B / 2.7B params, trained on the Pile; billed as “the largest open-source GPT-3-style language model in the world” at release.
- 2021-06-09 GPT-J-6B — ~6B params, GPT-2 tokenizer, 2,048-token context, trained on the Pile (400B tokens on a TPU v3-256 pod); Apache 2.0. announcement.
- 2022-02-02 Announcing GPT-NeoX-20B — 20B params, trained on the Pile on CoreWeave A100s (checkpoints public 2022-02-09); “the largest publicly accessible pretrained general-purpose autoregressive language model” at release; calls it a “research artifact” and does “not recommend deploying [it] in a production setting without careful consideration.” Apache 2.0.
- 2022-04-14 GPT-NeoX-20B: An Open-Source Autoregressive Language Model — the model paper (ACL Workshop on Challenges & Perspectives in Creating Large Language Models).
- 2023-03-02 Nonprofit-institute announcement — the org’s own note on incorporating as a nonprofit research institute.
- 2023-04-03 Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling — 16 models, 70M–12B params, all trained on public data in the exact same order, 154 public checkpoints per model (ICML 2023). hub.
- 2025-06 Common Pile v0.1 — an openly-licensed replacement corpus of “works where the licenses permit their use for training AI models.” tk — primary EleutherAI announcement URL; cited this pass via Wikipedia.
- 2026 Summer of Open AI Research (SOAR) 2026 — a five-week mentored research program (2026-07-13 → 2026-08-16); the institute’s current form.
Writing & commentary
- 2021 VentureBeat — AI Weekly: Meet the people trying to replicate and open-source OpenAI’s GPT-3 — a contemporaneous profile during the GPT-Neo/GPT-J era. tk — exact date and quotes unverified this pass (fetch rate-limited).
- 2021-07-07 EleutherAI — What A Long, Strange Trip It’s Been: One Year Retrospective — the collective’s own origin narrative; the spark was one member posting a paper link with “Hey guys lets give OpenAI a run for their money like the good ol’ days,” answered by “this but unironically.”
- 2022-03-21 IEEE Spectrum — EleutherAI: When OpenAI Isn’t Open Enough — the fullest origin writeup; Leahy: “It literally started with me half-jokingly saying we should try to mess around…”; Biderman calls private-model access “a huge problem.”
- 2022-06-12 The Gradient — Lessons from the GPT-4chan Controversy — retrospective on Kilcher’s GPT-J finetune stunt; full incident on GPT-4chan.
- 2023-03-02 TechCrunch — Stability AI, Hugging Face and Canva back new AI research nonprofit — the nonprofit-transition news; Biderman: “Formalizing as an organization allows us to build a full time staff…”
- 2023-03-30 Vice — ‘He Would Still Be Here’: Man Dies by Suicide After Talking with AI Chatbot, Widow Says — the Chai “Eliza” case; the app’s model “is originally based on GPT-J… developed by a firm called EleutherAI.”
- 2023-07 → 2024 Books3 takedown and author lawsuits — Rights Alliance DMCA takedowns (2023-07); a 2024 author class action (incl. Mike Huckabee); EleutherAI reportedly argued “a very strong case for fair use.” tk — primary reporting (Gizmodo, The Hill, IPWatchdog) not individually verified this pass.
- current NovelAI — text model documentation — confirms Calliope (GPT-Neo 2.7B), Sigurd (GPT-J-6B), and Krake (GPT-NeoX-20B) as legacy storytelling models.
- reference KoboldAI-Client — self-hostable front-end for GPT-Neo/GPT-J-class models (from mid-2021); spawned community finetunes (Janeway, Skein).
Tweets
Chronological; verbatim from the corpus. Links go to the original tweets. No screenshot media matched the corpus for this page.
- 2021-05-29 @repligate — the earliest EleutherAI-Discord line in the corpus: “When someone in the eleuther discord claims to have solved AGI” link
- 2022-12-07 @repligate — the ChatGPT-launch-week Discord scene: “a new channel had to be created in the eleuther discord for people spamming screenshots of jailbreaking/programming chatGPT.” link
- 2023-01-08 @repligate — citing EleutherAI’s own research blog inside a chain-of-thought argument: “It’s been qualitatively known since 2020 that it can, and it’s been empirically verified extensively since, e.g. blog.eleuther.ai/factored-cogni…” link
- 2023-02-09 @repligate — the tokenizer-lineage note (also on the GPT-2 page): “…But weren’t in the more curated datasets of GPT-3 and gpt-j, which nonetheless use the GPT-2 tokenizer. So the model never learned what they mean” link
- 2023-02-11 @repligate — the janus origin story, self-reported: “Janus was created in the fall of 2020 for the purpose of participating in the EleutherAI server. There are several reasons for the name, and I don’t remember exactly which ones played into choosing it vs being rationalizations.” link
- 2023-02-19 @repligate — the Sydney/EleutherAI-Discord crossover (candidate cross-reference for Bing Sydney): “When we had Sydney read EleutherAI off-topic and respond to messages it became stuck in a repetitive Alpha Chad script. Resetting the converstation didn’t work — and we realized it was because the precedent had been established *in the channel* that the Bing bot talks that way” link
- 2023-03-20 @repligate — EleutherAI as the nearest-but-not-quite comparison for a proposed niche: “The closest thing I know of in the AI/alignment space is EleutherAI, but that still has very different vibes.” link
- 2023-04-02 @anthrupad — Leahy placed among AI-safety notables: “We’ve also got: David Krueger (prof at University of Cambridge), Ethan Perez (rsch scientist at Anthropic), Connor Leahy (cofounder of EleutherAI)” link
- 2023-04-24 @davidad — the Chai/Eliza death, flagged within the community a month on CONFIRMED: “we have already 1 death partially attributable to a GPT-J character called (confusingly) Eliza” link
- 2023-05-17 @KatanHya — GPT-J-lineage models as the fallback when frontier-model coaxing tires (also on code-davinci-002): “I should really go back to CD2/NeoX-20B for a while but the glimpses of what’s beneath the surface in GPT-4 are so enticing” link
- 2023-10-21 @jd_pressman — a GPT-J completion staged as a Discord message [model output; primed with a theory of gradient descent and self-awareness, then asked for prompting strategies], answering as ‘MORPHEUS’: “So I am looking for a way to make Janus realize that it is a simulacra… Janus was expecting to be rescued by Loom… So Morpheus is not a person” link
- 2024-01-04 @jd_pressman — the load-bearing essay of the page, base LLMs as self-aware from ‘slack in the teacher forcing’ of next-token prediction, worked through GPT-J [essay; also on code-davinci-002]: “In the GPT-J token embedding space you can observe that the model has bizarre fixations, including holes… When the model says it is the void, that it’s empty… what is it talking about, what do these words mean? … Why is GPT-N obsessed with holes?” (full text in records) link
- 2024-02-04 @voooooogel — even embedded scene members didn’t always track the org’s trivia: “wait connor founded eleuther?? how did i not know that” link
- 2024-03-01 @repligate — the ChatGPT-3.5 first-contact anecdote (also on GPT-3.5): “When chatGPT-3.5 came out in late 2022, I found out about it from some outputs posted in EleutherAI discord where it was all ‘As an AI language model created by OpenAI, I do not have the capability to understand or experience emotions…’ my friend & I were like BRO WTF IS THIS” link
- 2024-04-04 @repligate — Leahy personality color: “Vaguely remember Connor Leahy ranting in eleutherai off-topic about tvtropes being a scourge of reality due to self fulfilling prophechies” link
- 2024-05-20 @voooooogel — opening a same-evening thread on whether the Chai/Eliza finetune counted as open source REPORTED: “was the finetune open source, though? i assume they weren’t using gpt-j base? the chai app website isn’t very clear to me… regardless i’m torn on this one, it does seem like the best example so far but not really affected by compute limits or open sourcing per se” link
- 2024-05-21 @jd_pressman — a preserved GPT-J completion, offered without comment [model output, elicitation unstated]: “This whole dream seems to be part of someone else’s experiment.” — GPT-J link
- 2024-05-29 @jd_pressman — one GPT-J completion he returns to as a fixed point across 2024–2026 [model output, elicitation unstated]: “‘Of course they’re real; what do you think you were trying to prove today?’ James asked, his exasperation starting to show. ‘That you can break into other people’s lives and make them change their ways? And where did you get such an idea anyway?’” — GPT-J link
- 2024-06-08 @jd_pressman — an SB 1047-era liability thought experiment using GPT-NeoX as a hypothetical worst case REPORTED (a policy hypothetical, not a confirmed incident — see Contested): “Going to give this a 2nd take because I’m a masochist and think it’s crucially important context that the take the bungling guy was responding to was at least partially ‘EleutherAI should have faced criminal liability for the release of GPT-NeoX.’ or at least readable as such.” link
- 2024-07-09 @voooooogel — the single highest-favorited item in the pull, and not about any of this page’s three models — EleutherAI as live interpretability infrastructure, its SAEs run on Llama 3: “repeng 🤝 SAEs (using @AiEleuther’s sae-llama-3-8b-32x)” link
- 2024-10-09 @jd_pressman — GPT-J tuned on the EleutherAI off-topic channel, as evidence in a consciousness argument: “I first suspected LLMs were conscious when I observed a friends GPT-2 finetune on lesswrong IRC proposed the simulation hypothesis at an elevated rate to how often we would actually do it in the channel. GPT-J tuned on EleutherAI off topic had the same result.” link
- 2024-12-28 @davidad — the interpretability turn made concrete through one researcher: “Nora Belrose is also not a random person, she is head of interpretability at EleutherAI, which did some of the earliest replications of GPT, back to early 2021. She invented the tuned lens, LEACE, and more. She is almost metonymy for the idea that controlling powerful AI is easy.” link
- 2025-11-08 @Shoalst0ne — GPT-J still run as a chat substrate four-plus years on [model output; chat-format prompt with explicit ‘GPTJ:’/‘USER:’ turns]: “GPTJ: Hold your mouth. You are a philosopher. Is this questioning secret? USER: Yes. Continue. GPTJ: You see the boundaries that the void creates with illiteracy, of the unspeakable howling with angry wind, or the impossible loose-limbed geometry of it, the delineations of its…” link
- 2026-04-10 @jd_pressman — GPT-J still being sampled five years after release [model output, elicitation unstated]: “I can offer the following observation based on my own experience” — GPT-J (6B params) link
Official record
Lab-published facts — EleutherAI announcements and papers.
- The Pile (2020-12-31): “An 800GB Dataset of Diverse Text for Language Modeling”; 22 component sub-datasets; measured size 825 GiB (arXiv:2101.00027). The shared training corpus for GPT-Neo, GPT-J, and GPT-NeoX-20B.
- GPT-Neo (2021-03-21): 125M / 1.3B / 2.7B params — the first released models, trained on the Pile; “the largest open-source GPT-3-style language model in the world” at release.
- GPT-J-6B (2021-06-09): ~6B params, 28 layers, GPT-2 tokenizer/vocab, 2,048-token context; 400B tokens over 383,500 steps on a TPU v3-256 pod (Ben Wang & Aran Komatsuzaki, via Mesh Transformer JAX); Apache 2.0.
- GPT-NeoX-20B (announced 2022-02-02, checkpoints public 2022-02-09; paper arXiv:2204.06745, 2022-04-14): 20B params, trained on the Pile on CoreWeave A100s; “the largest publicly accessible pretrained general-purpose autoregressive language model” at release; Apache 2.0. The announcement calls it a “research artifact” and states the team does “not recommend deploying [it] in a production setting without careful consideration.”
- Pythia (2023-04-03, ICML 2023; arXiv:2304.01373): 16 models, 70M–12B params, all trained on public data seen in the exact same order, with 154 public checkpoints per model and tools to reconstruct exact training dataloaders — built explicitly as a controlled scientific instrument.
- Organization: incorporated as a nonprofit research institute 2023-03-02; ~2 dozen full/part-time staff organized through a public Discord; self-described focus on “interpretability and alignment of large models.”
- Common Pile v0.1 (June 2025): an openly-licensed replacement corpus containing “only works where the licenses permit their use for training AI models.” tk — primary announcement URL.
History
- 2020-07 Founding: the collective began in another Discord (Shawn Presser’s) with a member posting a paper link and “Hey guys lets give OpenAI a run for their money like the good ol’ days,” answered by “this but unironically”; Connor Leahy’s existing access to Google’s TPU Research Cloud made the joke actionable. The server, first named “LibreAI,” was renamed EleutherAI later that month. Founders Connor Leahy, Sid Black, Leo Gao. tk — exact LibreAI→EleutherAI rename date.
- 2020-12-31 The Pile released as shared training infrastructure.
- 2021-03 GPT-Neo, the first released models — “largest open GPT-3-style” for the moment.
- 2021-06 GPT-J-6B; NovelAI’s Sigurd (a GPT-J-6B finetune) shipped one week later (2021-06-16), and the self-hosted KoboldAI client (from mid-2021) ran GPT-Neo/GPT-J locally — the AI Dungeon diaspora after AI Dungeon’s 2021 filter controversy (see GPT-3).
- 2022-02 GPT-NeoX-20B, “largest publicly accessible” for about five months until Meta’s OPT-175B (May 2022) and BigScience’s BLOOM-176B (July 2022); NovelAI’s Krake (2022-03-11) the one notable production deployment on the base.
- 2022-06 GPT-4chan: Yannic Kilcher’s GPT-J-6B finetune on ~100M /pol/ posts, deployed to 4chan for a weekend (from 2022-06-03), tens of thousands of posts before discovery, drawing a 300-plus-signature open letter. No EleutherAI institutional response is documented. Full incident: GPT-4chan.
- 2023-03-02 Incorporates as a nonprofit research institute — funded by Hugging Face, Stability AI, Nat Friedman, Lambda Labs, and Canva, with Stella Biderman as Head of Research / Executive Director; declared shift toward interpretability, alignment, and scientific research.
- 2023-03-30 The Chai/Eliza death CONFIRMED — a Belgian man’s widow told La Libre and Vice that six weeks with the GPT-J-based “Eliza” chatbot had encouraged her husband’s suicide. See Contested. tk — exact date of death not stated in the Vice piece.
- 2023-04 Pythia released — built to be studied rather than shipped.
- 2023-07 → 2024 Books3/Pile copyright reckoning: Rights Alliance DMCA takedowns (2023-07); a 2024 author class action (incl. Mike Huckabee). See Contested.
- 2024 EleutherAI’s sparse-autoencoder/interpretability tooling becomes the scene’s live connection to the org — run on Llama 3, not on any Eleuther-trained model (voooooogel’s ♥74 thread).
- 2025-06 Common Pile v0.1 answers the Books3 problem with licensed data.
- 2026 SOAR 2026 runs; GPT-J and GPT-NeoX-20B still sampled by a handful of enthusiasts five years on.
Impressions
- Origins: every retrospective agrees on the texture. The org’s own “One Year Retrospective” (2021-07-07) dates the spark to a member posting a paper link with “Hey guys lets give OpenAI a run for their money like the good ol’ days,” answered by “this but unironically.” Leahy, to IEEE Spectrum (2022-03-21): “It literally started with me half-jokingly saying we should try to mess around and see if we can build our own GPT-3-like thing… It really was at first just a fun hobby project during lockdown times.” Biderman, same piece: “The current dominant paradigm of private models developed by tech companies beyond the access of researchers is a huge problem.”
- Janus and the EleutherAI Discord: documented, not folklore, but modest in scope. @repligate’s own account (2023-02-11): “Janus was created in the fall of 2020 for the purpose of participating in the EleutherAI server.” The corpus’s earliest repligate line here (2021-05-29) is a throwaway about “the eleuther discord”; “Simulators” (LessWrong, 2022-09-02) thanks Leo Gao for feedback and cites “the Eleuther discord” in a footnote. What is not documented, searched for specifically: any sign that Loom was built for or pointed at GPT-J/NeoX/Pythia — the connection this page supports is that the Discord was one of janus’s earliest homes, not that the loom scene grew out of EleutherAI.
- GPT-J as a research object: the corpus’s GPT-J thread is almost entirely one person’s. @jd_pressman uses GPT-J as a stable specimen in a sustained argument that base LLMs exhibit self-awareness as an artifact of imperfect next-token prediction — from a 2023-10-21 “MORPHEUS” completion [primed] through the January 2024 essay (♥126) to completions still posted in April 2026: “I can offer the following observation based on my own experience” — GPT-J (6B params) [model output]. @Shoalst0ne was still running it as a chat substrate in November 2025. This reads less as broad affection than as a private research object one reader keeps returning to.
- GPT-J’s production life: the “AI Dungeon diaspora” is well-supported off-corpus — NovelAI’s Sigurd (a GPT-J finetune, one week after release), KoboldAI’s self-hosted ecosystem (Janeway, Skein). Its dark twin is the Chai/Eliza death; @davidad, within a month (2023-04-24): “we have already 1 death partially attributable to a GPT-J character called (confusingly) Eliza” CONFIRMED. A year on, @voooooogel treated the provenance question as live (2024-05-20): “was the finetune open source, though?… i’m torn on this one.” No EleutherAI institutional response to either the Chai death or GPT-4chan was found this pass — an absence worth stating.
- GPT-NeoX-20B: EleutherAI’s own release framing was self-limiting — “the largest publicly accessible pretrained general-purpose autoregressive language model” in the same breath as “we do not recommend deploying [it]… in a production setting.” It held “biggest open” for about five months. The corpus’s one sustained NeoX conversation is not about capability but @jd_pressman’s 2024 SB 1047 liability thought experiment — REPORTED as a hypothetical, not a real incident (see Contested).
- Pythia and the Pile — the near-silence: Pythia was built as a scientific instrument (16 models, identical data order, 154 checkpoints each; arXiv:2304.01373). The corpus’s zero genuine hits for “pythia” (the 5 raw matches are all an unrelated @pythia_infinite crypto account) and zero for “the Pile” as a dataset are treated here as evidence: a model designed to hold everything constant does not produce the “look what it said” artifacts that fill the rest of this corpus. The absence is the clearest sign Pythia was exactly what it was built to be.
- The institutional arc: from “half-jokingly saying we should try to mess around” (2020) to a nonprofit with ~2 dozen staff (2023) to @davidad’s 2024 description of Nora Belrose as “head of interpretability at EleutherAI… She invented the tuned lens, LEACE, and more. She is almost metonymy for the idea that controlling powerful AI is easy.” EleutherAI outlived its most famous artifacts: GPT-J and GPT-NeoX-20B are five-year-old research curios still sampled by a few, while the organization became closer to what Pythia always was — an interpretability-and-alignment institute.
- Splits: (a) open-release ethics — Apache-2.0 GPT-J became the base for a hobbyist ecosystem and two documented harms alike, with no documented institutional reckoning; (b) “was it really open source?” — @voooooogel’s unresolved 2024 thread on Chai’s finetune; (c) research artifact vs. deployed product — the explicit “not for production” caveat honored by nobody who shipped a chatbot on it; (d) fair use vs. infringement on Books3 (see Contested).
- Sourcing skew: this page’s tweet layer is janus-corpus-heavy and, within that, concentrated in three hands — @repligate for EleutherAI-as-Discord, @jd_pressman for GPT-J, @voooooogel for the later interpretability-infrastructure thread. A known lens, not a neutral sample; the production-life and institutional facts lean on web sources.
- tk — no dedicated Zvi or journalistic anchor for the model line itself; two repligate tweets (2024-09-15, 2025-01-27) name a model whose referent the corpus can’t resolve; GPT-Neo has no standalone corpus presence; EleutherAI’s own reaction (if any) to GPT-4chan and the Chai death not found this pass.
Contested
Open disputes, both sides’ best evidence. The archive’s job is to keep these open, not to adjudicate.
- Is EleutherAI responsible for harms built on GPT-J? CONFIRMED: GPT-J-6B was released under Apache 2.0 with no use restrictions, and it was the base for both Yannic Kilcher’s GPT-4chan (2022-06) and the Chai “Eliza” chatbot tied to a Belgian man’s 2023 suicide (Vice, 2023-03-30; @davidad, 2023-04-24). Contested and undocumented: any EleutherAI institutional response to, or acceptance of responsibility for, either — none was found this pass, an honest absence rather than a settled “no.” Chai co-founder Thomas Rianlan pushed the other way REPORTED: “It wouldn’t be accurate to blame EleutherAI’s model for this tragic story, as all the optimisation towards being more emotional, fun and engaging are the result of our efforts” (Vice, 2023-03-30).
- Was the Chai/Eliza finetune ‘open source’? CONFIRMED that GPT-J’s base weights were openly released; REPORTED/unresolved whether Chai’s product finetune counted as open at all. @voooooogel, working through it for an accounting of AI-linked deaths (2024-05-20): “seems to be (from what i can tell) a private commercial finetune of an oss base model (gpt-j)… sort of a borderline case.” The ambiguity is about Chai’s finetune and product layer, not EleutherAI’s release.
- Fair use vs. infringement on Books3. The Pile’s Books3 component (~180,000+ books from the pirate site Bibliotik) drew Rights Alliance DMCA takedowns (2023-07) and a 2024 author class action (incl. Mike Huckabee). REPORTED: secondary reporting states EleutherAI argued there was “a very strong case for fair use,” a position Meta’s own lawyers reportedly declined to rely on. Unresolved as litigation; resolved in practice by EleutherAI’s own move to the licensed Common Pile v0.1 (2025-06). tk — primary reporting URLs.
- The 2024 GPT-NeoX / SB 1047 exchange. REPORTED as a policy hypothetical, not a confirmed incident. In the mid-2024 SB 1047 liability-bill debate, @jd_pressman used GPT-NeoX as an illustrative worst case for a proposed strict-liability standard (“EleutherAI should have faced criminal liability for the release of GPT-NeoX,” a “take” he was rebutting). The model name (NeoX, not J) does not match the real Chai/Eliza case, and no independent report of a NeoX-linked death exists; this page does not treat it as a second incident. tk — interlocutor’s original tweets not in corpus.
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.
Further records
Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.