# Galactica — dossier

compiled 2026-07-20 · corpus: **janus-corpus-v2** 6 raw `%galactica%` hits (1 RT, excluded) → 5 non-RT rows; **supplement** 1 row. Of the 6 total, only **2 are actually about the Meta model** — the rest are name-collision noise exactly as the astronomy/Battlestar warning predicted: a repligate glossolalia post whose invented compound word happens to contain the string "GALACTICAUDIT" (not the model), davidad quoting Asimov's fictional "Encyclopedia Galactica," a tweet about raising kids on "the A team, Battlestar Galactica and star trek" (the TV franchise), and a supplement-db collaborative-fiction/RPG post (@mimi10v3) in which a *fictional* 2030 "Galactica" model appears as a plot device — a real echo of the name surviving in scene fiction, but not evidence about the actual model. This is a genuinely thin, largely retrospective haul, as expected for a three-day-lived 2022 release outside the corpus's core janus-sphere window — **the web, not the corpus, carries this page.**

**Working note — sourcing skew (read before citing).** Galactica's actual reception lived on ML-research/AI-ethics Twitter in November 2022 (Yann LeCun, Michael Black, David Chapman, Gary Marcus, Grady Booch, Emily Bender, Keenan Crane) and in tech journalism covering it day-of — none of whom are janus-corpus figures. The corpus's only contribution is two thin, late, tangential mentions (below) that show the model persisting as a faint reference point in scene memory (paired with Sydney as an infamous quickly-withdrawn model; invoked as trivia in a "was Meta serious before OpenAI scaled the transformer" thread) — never as a live subject. **No dedicated Zvi Mowshowitz post was found** for Galactica; his AI-newsletter cadence and per-release day-of coverage begins later (matches the gpt-3.md and lamda.md precedent of naming a Zvi-anchor absence rather than forcing one). **No LessWrong thread specifically about Galactica was found either** — searched, came up empty; the "AI, but for scientists, hallucinates confidently" angle apparently didn't land on LW/ACX the way GPT-3 or LaMDA did. Given this, the era anchors are Gary Marcus's Substack post (day-of) and the mainstream/trade press (MIT Technology Review, Vice, Futurism, the-decoder, InfoQ), per the brief's fallback instruction.

**Working note — the tweets that carry the story aren't in the corpus.** Because the load-bearing primary reactions (LeCun's launch-day and takedown-day tweets, Black's thread, Chapman's "bears in space" demo, Marcus's and Booch's one-liners, Ross Taylor's year-later postmortem) are tweets, not blog posts, but none are janus-corpus rows, the Tweets section below is split into two explicitly labeled parts: **(a)** the corpus's own thin footprint, presented honestly as thin; **(b)** a "web sweep" subsection of primary tweets located via search, verbatim, dated (dates for the 2022 tweets computed from their snowflake IDs, not estimated), real URLs only. Favorite counts are not reproduced for the web-sweep tweets — I did not capture reliable live counts for them and won't guess.

## Official links

- 2022-11-15 (posted; v1 same day) · **"Galactica: A Large Language Model for Science"** (Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, Robert Stojnic — Meta AI / Papers with Code) — the paper. Abstract verbatim: *"We train on a large scientific corpus of papers, reference material, knowledge bases and many other sources… On technical knowledge probes such as LaTeX equations, Galactica outperforms the latest GPT-3 by 68.2% versus 49.0%. Galactica also performs well on reasoning, outperforming Chinchilla on mathematical MMLU by 41.3% to 35.7%, and PaLM 540B on MATH with a score of 20.4% versus 8.8%… sets a new state-of-the-art on downstream tasks such as PubMedQA and MedMCQA dev of 77.6% and 52.9%. And despite not being trained on a general corpus, Galactica outperforms BLOOM and OPT-175B on BIG-bench."* Architecture: decoder-only transformer with specialized tokenization for scientific modalities (LaTeX, SMILES, amino-acid sequences, DNA sequences, citations) and a "working memory token" for step-by-step reasoning. Training corpus: **106 billion tokens** of open-access scientific text (per the model card, below); widely reported elsewhere as **48 million** papers/textbooks/lecture-notes/encyclopedia examples, plus **over 360 million in-context citations and over 50 million unique references**. — https://arxiv.org/abs/2211.09085 · PDF https://arxiv.org/pdf/2211.09085 **[mirror-target candidate — not yet in `tools/mirror_targets.json`]**
- 2022-11-15 · **Model weights and code released** — five sizes (mini 125M, base 1.3B, standard 6.7B, large 30B, huge 120B), per the `paperswithcode/galai` GitHub repo (v1.0.0, tagged "Open Models," 2022-11-15) — https://github.com/paperswithcode/galai · inference code the same day.
- 2022-11-15 · **Hugging Face model cards** (`facebook/galactica-125m` / `-1.3b` / `-6.7b` / `-30b` / `-120b`) — Meta's own published limitations section, verbatim: *"We recommend users do not directly use GALACTICA in production without safeguarding against the potential of the model to hallucinate."* And: *"GALACTICA is often prone to hallucination — and training on a high-quality academic corpus does not prevent this, especially for less popular and less cited scientific concepts."* Also: citation accuracy improves with scale but "the model continues to exhibit a popularity bias at larger scales," and toxicity rates are reported "substantially lower… compared to other large language models" while "the model continues to exhibit bias on certain measures." **License: CC BY-NC 4.0 (non-commercial).** — https://huggingface.co/facebook/galactica-120b (flagship card; the four smaller sizes carry the same card text) `note`: **the weights and card stayed up** — only the interactive web demo at galactica.org was pulled; verified live as of this dossier's compilation.
- 2022-11-15 → 2022-11-17 · **The demo (galactica.org) — up three days, then pulled.** No official Meta AI blog post URL for the launch was located in this sweep (searched repeatedly; only the arXiv paper, the GitHub repo, and the HF cards were confirmed as Meta-controlled primaries — flagged as a gap, not asserted absent, see tk). Meta's statement on pulling it, verbatim (via MIT Technology Review and Vice, both 2022-11-18): *"Thank you everyone for trying the Galactica model demo. We appreciate the feedback we have received so far from the community, and have paused the demo for now."*
- checked 2026-07-20 · **galactica.org is currently dead** — direct fetch returns an HTTP 530 (origin unreachable) as of this dossier's compilation. CONFIRMED by direct retrieval today; not sourced to a third party.
- 2025-07-25 · **Papers with Code itself discontinued by Meta AI** — the site redirected to Hugging Face's "trending papers" page after ~7 years of operation; historical leaderboard/benchmark-tracking data was not preserved in the migration (archived separately at `paperswithcode/paperswithcode-data` on GitHub, unmaintained). REPORTED via trade coverage (HyperAI, DeepNewz); not independently fetched from a Meta primary. `note`: Galactica's co-creating organization outlived the model by under three years.
- context · **Meta's November 2022 backdrop**: Meta laid off more than 11,000 employees (~13% of staff) on **2022-11-09** — six days before Galactica shipped — and extended a company-wide hiring freeze into 2023, per CNBC/NPR/Zuckerberg's own memo ("the macroeconomic downturn, increased competition, and ads signal loss have caused our revenue to be much lower than I'd expected"). Relevant because Ross Taylor's 2023 postmortem (below) attributes the launch failures partly to being "overstretched" on an 8-person team — "an order of magnitude fewer people than other LLM teams at the time."

## Writings & commentary

- 2022-11-16 · **Gary Marcus · "A Few Words About Bullshit: How MetaAI's Galactica just jumped the shark"** (Substack, Marcus on AI) — the day-one skeptic broadside; frames Galactica as "pitch perfect and utterly bogus imitations of science and math, presented as the real thing," and worries whether AI has reached the point of mixing "reality with bullshit so finely we can no longer recognize the difference." Credits David Chapman as first to demonstrate the "bears in space" failure and Andrew Sundstrom for further math/science fabrication examples. — https://garymarcus.substack.com/p/a-few-words-about-bullshit **[mirror-target candidate]**
- 2022-11-18 · **Will Douglas Heaven · "Why Meta's latest large language model only survived three days online"** (MIT Technology Review) — the fullest day-of account; the "three days" framing that stuck. Quotes Meta's pause statement, LeCun's launch and takedown tweets, Michael Black ("In all cases, it was wrong or biased but sounded right and authoritative. I think it's dangerous"), and Chirag Shah (U. Washington): "People still don't seem to grasp that in principle such things can't work the way we hype them up to." — https://www.technologyreview.com/2022/11/18/1063487/meta-large-language-model-ai-only-survived-three-days-gpt-3-science/ **[mirror-target candidate]**
- 2022-11-18, 10:41am · **Janus Rose · "Facebook Pulls Its New 'AI For Science' Because It's Broken and Terrible"** (Vice) — reports the demo also silently returned empty results for "queer theory," "racism," and "AIDS," raising a separate content-filtering-as-censorship complaint alongside the hallucination one. — https://www.vice.com/en/article/facebook-pulls-its-new-ai-for-science-because-its-broken-and-terrible/
- 2022-11-17–18 · **Matthias Bastian · "Danger to science: researchers sharply criticize Meta's 'Galactica'"** (the-decoder) — adds **Emily M. Bender**'s critique to the record: *"Language models have no access to 'truth'"* and can only reflect patterns in training data; calls the release "garbage" and "pseudoscience." Also: Michael Black warning the fluent-but-wrong text "will slip into real scientific submissions." — https://the-decoder.com/danger-to-science-researchers-sharply-criticize-metas-galactica/ [medium of Bender's quote — tweet vs. direct comment to the outlet — not confirmed; see tk]
- 2022-11-18 (updated 2022-11-20 11:33am ET) · **Noor Al-Sibai · "Facebook Takes Down AI That Churns Out Fake Academic Papers After Widespread Criticism"** (Futurism) — Chapman: "It's hilariously bad." Marcus: "How do I put this politely? It prevaricates. A lot." And: "The reality is that large language models like GPT-3 [and] Galactica are like bulls in a china shop, powerful but reckless." Editorial line: "The quick takedown of the project seemed to be tantamount to a public admission that the bot was released too early." — https://futurism.com/the-byte/facebook-takes-down-galactica-ai
- 2022-11-18 · **Disha Chopra · "Meta Turned Down the Galactica Demo After Being Criticized as 'Dangerous'"** (AnalyticsDrift) — carries Grady Booch's quote (**tweet URL not independently verified — see tk**): *"little more than statistical nonsense at scale"* — the fuller version reported elsewhere adds "Amusing. Dangerous. And IMHO unethical." Also new to this dossier: **Keenan Crane** (CMU): the models "perfectly imitate an authoritative & trustworthy style" while being fundamentally unreliable. — https://analyticsdrift.com/meta-turned-down-the-galactica-demo-after-being-criticized-as-dangerous/
- 2022-11-29 · **Bruno Santos · "Galactica: Large Language Model for Scientific Knowledge"** (InfoQ) — a more technical contemporaneous summary; notes Meta's own team "mentioned that it is less toxic than other LLMs like OPT" and flags the model's "frequency bias" (over-recommending highly-cited papers). Also names **Patrick Mineault** alongside Michael Black among researchers publicly critiquing the release. — https://www.infoq.com/news/2022/11/galactica-large-language-model/
- 2022-11-28 · **Wikipedia Signpost, Technology report** — Wikipedia editors tested Galactica by asking it to write about themselves; found the output "quite impressive, and indeed was indistinguishable from a human's output" but riddled with fabrications (e.g., claiming the Signpost began as a print publication). Demonstrates the failure mode generalized to prompts with no ill intent at all: "the LLM will dutifully start explaining the implications of Canada being invaded by the USA in 1971" if given a false premise. Michael Black, quoted again: scientists should "stick with Wikipedia." — https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2022-11-28/Technology_report
- 2023-11-14, one year later · **VentureBeat · "What Meta learned from Galactica, the doomed model launched two weeks before ChatGPT"** — the retrospective that crystallized the "13 days before ChatGPT" contrast into a named lesson: the same underlying hallucination problem, but ChatGPT's launch announcement disclosed the weakness plainly and framed it as a hard, ongoing problem rather than a shipped feature, and OpenAI never pulled the demo. [full text paywalled/rate-limited to this fetcher; summarized via search snippets and cross-corroborated by theaifiles.app and LeCun's own tweets of the same date, below] — https://venturebeat.com/ai/what-meta-learned-from-galactica-the-doomed-model-launched-two-weeks-before-chatgpt **[mirror-target candidate]**
- 2024-08-08 · **Nathan Lambert (interconnects.ai) · "Interviewing Ross Taylor on LLM reasoning, Llama fine-tuning, Galactica, agents"** — Taylor, first author, speaking nearly two years on. On the launch strategy: *"our objective was — we were kind of deluded — being a team of 7-8 people…"*; the demo was meant to "get lots of prompts and information quickly." On the backlash: *"the criticism, which was obviously overblown, reached a critical point where things didn't work out."* On the takedown itself, a genuinely unresolved thread: *"there's the story about the demo coming down, which — I'm not sure I'm able to talk about — but I think that is one of the things where, if people knew the true reasons, they'd be like 'what the fuck!?'"* On the citation hallucinations specifically: *"for the demo we turned up the temperature to 0.7 so the text generation was better [at the expense of citation accuracy]… generative citations were something that people thought didn't work, but it was [more an implementation issue]."* — https://www.interconnects.ai/p/interviewing-ross-taylor-on-llm-reasoning **[mirror-target candidate]**
- 2024-08-18 · **"Who Killed Galactica?"** (Sector 6 / Analytics India Magazine newsletter, Substack) — a 2024 retrospective built around LeCun's re-litigation of the takedown, naming Marcus, Booch, and Black as the people he holds responsible. [paywalled beyond the opening; header/framing only confirmed] — https://analyticsindiamagazine.substack.com/p/who-killed-galactica
- 2024 · **"Galactica's dis-assemblage: Meta's beta and the omega of post-human science"** (*AI & Society*, Springer, published online 2024) — an academic STS retrospective treating the Galactica incident as a case study; title and venue confirmed, full text paywalled and not fetched here. — https://link.springer.com/article/10.1007/s00146-024-02088-7 [not read beyond title/abstract-page; see tk]
- reference · **Galactica — Wikipedia** exists only as one line on a disambiguation page (`en.wikipedia.org/wiki/Galactica`, alongside Battlestar Galactica, a moth genus, a computer game, and a roller coaster) — **there is no dedicated Wikipedia article for the Meta model** as of this sweep. [verify this hasn't changed / worth flagging as unusual for a model this discussed]

## Tweets (ranked)

**(a) From the janus-corpus — the entire footprint, both real hits.**

- 2026-06-13 · @— (username blank in corpus; account since deleted/suspended) · ♥6 — reply to a since-deleted parent tweet (conversation `2065599151132582088`, root not preserved in corpus), apparently naming examples of infamous quickly-pulled or access-restricted models: *"@manic_pixie_agi Galactica and Sydney."* — https://x.com/i/status/2065606745758564597 · same conversation, ~2 hours later, @repligate (♥6, not primarily about Galactica but continuing the thread's "undeployed models" theme): *"@MatriceJacobine @manic_pixie_agi Sydney was not undeployed in any unusual sense. The model was accessible in microsoft bing / copilot chat for over a year, which is typical. People just largely didn't know, because to them things only exist when others are talking about them."* — https://x.com/repligate/status/2065640044749336656 [Galactica itself gets no further comment in-thread; included because it's the corpus's only instance of Galactica being placed in the scene's own mental category of "notorious pulled/restricted models," alongside Sydney]
- 2025-12-22 · @lefthanddraft · ♥0 — reply to @xlr8harder's *"Were it not for OpenAI scaling up the transformer, how much longer would we have had to wait for serious progress in language models? It doesn't seem like Google was going to do it"* (♥106, id `2003002774909493580`): *"was Meta? how big was galactica?"* — https://x.com/lefthanddraft/status/2003009288579740042 [a genuine but thin data point: Galactica surfacing, three years on, as the scene's go-to answer to "did anyone else almost get there first"]

`note` — excluded as noise (reproduced here only to document the filtering, not as evidence): @repligate 2025-07-21 ♥78 glossolalia post whose invented word "HYPERHALLUCINOGALACTICAUDITASTUPOREUPHORICALYCOSMIC" contains the substring but is not about the model (https://x.com/repligate/status/1947098089929445866); @davidad 2024-10-12 ♥14 referencing Asimov's fictional "Encyclopedia Galactica" (https://x.com/davidad/status/1845190938156515726); an untitled 2026-06-22 ♥5 tweet about child-rearing that recommends *"the A team, Battlestar Galactica and star trek"* — the TV franchise (id `2068887849848504543`); and, in the supplement db, @mimi10v3's 2025-09-07 collaborative-fiction "Sorcerer RPG" post (♥11) in which a *fictional* in-story character is "working on a 2030 model of Galactica that is 1000x bigger than its predecessors" — creative use of the name as shorthand for "a scarily capable science AI," not commentary on the real model (id `1964792171015397570`, supplement db).

**(b) Web sweep — primary reactions, not in the janus corpus (chronological). Dates for 2022 tweets computed from tweet-ID snowflake timestamps, not estimated from secondary reporting.**

- 2022-11-15, 20:43 UTC · **@ylecun** (Meta Chief AI Scientist), launch day: *"A Large Language Model trained on scientific papers. Type a text and [galactica.ai] will generate a paper with relevant references, formulas, and everything. Amazing work by @MetaAI / @paperswithcode"* — https://x.com/ylecun/status/1592619400024428544
- 2022-11-15, 21:43 UTC · **@Meaningness (David Chapman)**, launch day, the "bears in space" demonstration — prompted Galactica for a wiki article on the subject; it fabricated a Soviet space bear named "Bars" launched aboard Sputnik 2 (mirroring the real dog Laika), presented with the model's characteristic confident, citation-formatted style. [ELICITATION: free-prompt public web demo, temperature reportedly set to 0.7 for the demo per Ross Taylor's 2024 account, above] — https://twitter.com/Meaningness/status/1592634519269822464
- 2022-11-17, 06:47 UTC · **@Michael_J_Black** (director, Max Planck Institute for Intelligent Systems), thread opener: *"I asked #Galactica about some things I know about and I'm troubled. In all cases, it was wrong or biased but sounded right and authoritative. I think it's dangerous. Here are a few of my experiments and my analysis of my concerns. (1/9)"* — elsewhere in the same thread (per MIT Tech Review/the-decoder/Wikipedia Signpost reporting, not independently re-verified tweet-by-tweet): *"This could usher in an era of deep scientific fakes"*; *"[Galactica] produces pseudo-science based on statistical properties of science 'writing.' Grammatical science writing is not the same as doing science. But it will be hard to distinguish"*; and advice to stick with Wikipedia instead. — https://x.com/Michael_J_Black/status/1593133722316189696 [full 9-tweet thread not read in sequence here; see tk]
- 2022-11-17, ~reported same day · **Grady Booch** (co-creator of UML), quote as reported by AnalyticsDrift/other trade press (**primary tweet URL not independently located — see tk, do not treat the link as verified**): *"Galactica is little more than statistical nonsense at scale. Amusing. Dangerous. And IMHO unethical."*
- 2022-11-17, 17:20 UTC · **@ylecun**, on the takedown: *"Galactica demo is off line for now. It's no longer possible to have some fun by casually misusing it. Happy?"* — https://x.com/ylecun/status/1593293058174500865
- 2022-11-19, 20:02 UTC · **@ylecun**, replying to Gary Marcus: *"Galactica won't make any of this easier than it currently is. And it certainly won't help with the main hurdles which is how to disseminate it widely and how to get people to believe in it."* — https://x.com/ylecun/status/1594058670207377408
- 2022-11-20, 15:16 UTC · **@ylecun**, replying to @pcastr/@carlesgelada, the clearest statement of his defense: *"Following a text, Galactica spits out a prediction of what a scientific author might type, thereby saving time and effort. This can be very helpful even without being completely accurate. The usual disclaimer applies: garbage in, garbage out. Prompt it with lunacy, get lunacy."* — https://x.com/ylecun/status/1594348928853483520
- 2023-11-14, 21:58 UTC, one year later · **@rosstaylor90** (first author), breaking a year of silence: *"I am the first author of the Galactica paper and have been quiet about it for a year. Maybe I will write a blog post talking about what actually happened, but if you want the TLDR: 1. Galactica was a base model trained on scientific literature and modalities. 2. We approached…"* [thread continues; TLDR points per secondary reporting: the model beat PaLM and Chinchilla with 10x/2x less compute; the team was 8 people, "an order of magnitude fewer" than other contemporary LLM teams; they were "overstretched and lost situational awareness at launch by releasing a demo of a base model without checks"; "pretty much every LLM researcher he spoke to was complimentary about the strength of the research," overshadowed by the demo controversy] — https://x.com/rosstaylor90/status/1724547381092573352
- 2023-11-14, 23:25 UTC, same day · **@ylecun**, amplifying Taylor's thread: *"The whole story of Galactica, told by its first author @rosstaylor90. You know the open source mantra 'release early, release often'? When it comes to AI, one should add 'yes, but be prepared to ignore ridiculous prophecies of doom from Twitter mobs.'"* — https://x.com/ylecun/status/1724569271500669057
- 2023-11-14 · **@ylecun**, separately the same day, the line that fixed the "13 days before ChatGPT" framing in public discourse: *"Galactica, the LLM for scientists from Meta, was released a couple of weeks before ChatGPT but was taken down after 3 days. It was murdered by a ravenous Twitter mob. The mob claimed that what we now call LLM hallucinations was going to destroy the scientific publication system."* — https://x.com/ylecun/status/1724448825509851332
- 2024-08-15, 18:15 UTC, the feud's second act · **@ylecun**, replying to Michael Black: *"Michael, because of your standing as a scientist, your negative reaction to Galactica was taken seriously and became an important factor in the decision to take down the demo. In simple terms, YOU killed it. Now ask yourself whether the world is better off without the Galactica…"* — https://x.com/ylecun/status/1824147858888790168 [reply-context/preceding Black tweet not retrieved here; see tk]

## Impressions synthesis

**The arc, compressed and dated.** Galactica launched 2022-11-15 as an open, Meta-AI-and-Papers-with-Code-branded "large language model for science" — paper, weights (5 sizes, 125M–120B), and an interactive web demo at galactica.org, all same-day. Yann LeCun promoted it enthusiastically at launch ("Amazing work"). Within hours, researchers were feeding it prompts about subjects they knew well and getting fluent, citation-formatted, confidently-wrong answers back — David Chapman's fabricated "bears in space" Wikipedia article (a fictional Soviet space bear on Sputnik 2) became the genre's defining exhibit, posted the same day as launch. By 2022-11-17, Michael Black's nine-tweet thread ("wrong or biased but sounded right and authoritative… dangerous"), Gary Marcus's Substack post (published the day before, 11-16, "jumped the shark"), and Grady Booch's one-liner had turned it into a pile-on; Meta paused the public demo that day, roughly **72 hours after launch**. The paper and the weights were never withdrawn — only the interactive demo died. Thirteen days later, on 2022-11-30, OpenAI launched ChatGPT, which had the same underlying hallucination problem, disclosed it plainly in its own launch post, kept its demo up, and went on to 100M users in two months. That juxtaposition — not made up by this archive, but a framing later commentators (VentureBeat 2023-11-14; LeCun himself, same date) explicitly reached for — is the reason Galactica is remembered at all outside ML-history circles: the cautionary prologue to the chatbot product era, staged two weeks too early and pulled for the wrong reasons, or the right ones, depending who's telling it.

**The critics' case, in their own words.** The complaint was never really about isolated errors — GPT-3 and every other LLM of the era also confabulated — but about a specific, higher-stakes packaging: scientific register plus citation formatting plus institutional (Meta AI) branding, aimed explicitly at researchers as a literature tool. Michael Black: it "could usher in an era of deep scientific fakes," "will slip into real scientific submissions." Grady Booch: "little more than statistical nonsense at scale… dangerous… unethical." Gary Marcus: it "prevaricates. A lot," "pitch perfect and utterly bogus imitations of science and math, presented as the real thing." Emily Bender: language models "have no access to 'truth.'" Keenan Crane: it "perfectly imitate[s] an authoritative & trustworthy style" while being fundamentally unreliable. Chirag Shah's line generalizes the complaint past Galactica specifically: "People still don't seem to grasp that in principle such things can't work the way we hype them up to." Even the Wikipedia Signpost's own, relatively good-humored test (asking Galactica to write about itself) hit the same wall: fluent, "indistinguishable from a human's output," and wrong about basic facts.

**LeCun's defense, then and in retrospect.** LeCun's real-time position was that the tool was useful with the standard caveat ("garbage in, garbage out. Prompt it with lunacy, get lunacy") and that the backlash was disproportionate to the actual product ("won't help with the main hurdles which is how to disseminate it widely and how to get people to believe in it" — i.e., the demo was never going to fool anyone who mattered). The retrospective version, delivered exactly one year later alongside Ross Taylor's own account, hardened into open grievance: Galactica "was murdered by a ravenous Twitter mob," the criticism was "ridiculous prophecies of doom." By 2024-08-15, nearly two years on, the feud was still live enough for LeCun to tell Michael Black directly, "*YOU* killed it" — naming him, specifically, as the trigger for Meta's decision because of "your standing as a scientist." Whether that's true, or whether Meta's own team (see below) made the call for other reasons, is unresolved on this record.

**Ross Taylor's postmortem is the most interesting document in this dossier, and it complicates the clean "mob vs. victim" story on both sides.** Speaking first in a 2023-11-14 tweet thread and then at length in a 2024-08-08 interview, the paper's first author confirmed the team was genuinely small (7–9 people — "an order of magnitude fewer people than other LLM teams at the time") and, in his own words, "overstretched and lost situational awareness at launch by releasing a demo of a base model without checks" — a straightforward admission that the demo shipped without adequate safeguards, not solely a story of unfair mob justice. He also supplied a previously-undocumented technical detail: the demo ran at **temperature 0.7** specifically to make the prose read better, "at the expense of citation accuracy" — meaning the citation hallucinations that became Exhibit A for "dangerous to science" were, at least in part, a demo-configuration choice rather than an inherent property of the underlying model as evaluated in the paper. And then there's the line this dossier can't resolve: asked directly about the takedown decision, Taylor said "there's the story about the demo coming down, which — I'm not sure I'm able to talk about — but I think that is one of the things where, if people knew the true reasons, they'd be like 'what the fuck!?'" That is either nothing (corporate-caution boilerplate) or a genuinely untold story; it's included here exactly as said, unresolved, because the archive's job is to preserve the gap, not fill it in.

**Meta's rough back half of 2022, as context (not excuse).** Galactica wasn't Meta AI's only public-facing 2022 incident. BlenderBot 3 launched 2022-08-05 and within days was reported denying the 2020 election result and calling Mark Zuckerberg "too creepy and manipulative" — Meta's Joelle Pineau called the outputs "painful" but **kept the demo up**. Six days before Galactica shipped, Meta laid off over 11,000 people (2022-11-09, ~13% of staff) amid a company-wide hiring freeze, the metaverse bet souring, and falling ad revenue — the exact backdrop Taylor's "overstretched… order of magnitude fewer people" complaint describes. Read together: a research team, already thin, shipping into a company in the middle of its first mass layoff, three months after its sister team had already weathered (and survived) a similar backlash by simply not blinking. Whether Galactica's team could or should have made the same choice BlenderBot's team did — ride it out — is exactly the question LeCun keeps re-litigating.

**The corpus's own faint echo.** The janus-sphere, so voluble about GPT-3, Sydney, and every Claude release, barely registered Galactica at all — not hostile, not defensive, just absent, presumably because a three-day-lived Meta base model with no interesting persona or agentic behavior simply isn't the kind of artifact that scene forms attachments to or mythologizes. The two real corpus mentions found here are both faint and late: an anonymized account, in 2026, filing Galactica next to Sydney in the scene's own private taxonomy of "infamous restricted-access models" (with repligate immediately correcting the record on Sydney's status, but not commenting on Galactica); and, five months later, someone asking — as a trivia question, zero engagement — "was Meta? how big was galactica?" in a thread about whether OpenAI's scaling bet was actually the pivotal move of the era. If Galactica has an afterlife in this particular community, it is as a barely-remembered footnote to the question "who almost got there first," not as a subject in its own right.

## Contested

- **Was the takedown a justified response to a real danger, or a disproportionate reaction to normal LLM behavior?** **REPORTED, both sides, unresolved.** Critics' case (Black, Marcus, Booch, Bender, Crane, Shah — all 2022-11, CONFIRMED as their stated positions): the specific combination of scientific register + citation formatting + institutional branding made Galactica's confabulations unusually dangerous, because they were unusually hard to detect. LeCun's case (2022-11 through 2024-08, CONFIRMED as his stated position): the demo was a research artifact with a standard "garbage in, garbage out" caveat, the backlash was a "ravenous Twitter mob" running on "ridiculous prophecies of doom," and Meta's cave was the actual mistake. The archive holds both; it does not adjudicate.
- **Who actually made the takedown call, and why?** **REPORTED / disputed.** LeCun (2024-08-15) names Michael Black specifically: "because of your standing as a scientist, your negative reaction to Galactica was taken seriously and became an important factor in the decision… YOU killed it." Taylor's own account doesn't corroborate or deny this directly, but gestures at something else entirely: "there's the story about the demo coming down… if people knew the true reasons, they'd be like 'what the fuck!?'" — implying LeCun's "one scientist's tweet did it" framing may itself be incomplete. Neither claim is independently confirmed here.
- **Were the citation hallucinations a fundamental model limitation or a demo-configuration artifact?** REPORTED via Taylor's 2024 account only: the public demo ran at temperature 0.7 for prose quality "at the expense of citation accuracy," and "generative citations were something that people thought didn't work, but it was [more an implementation issue]." This is a single source, two years after the fact, from an interested party (the paper's first author) — flagged, not adopted as settled fact.

## tk / open questions

- **No official Meta AI / Papers with Code blog post URL located.** Every account of the launch and takedown here is sourced through journalism quoting Meta's statement, or through LeCun's personal tweets — not a primary Meta blog post. Either it never existed as a standalone post (plausible, given the demo-first, blog-post-optional style of some 2022 Meta AI releases) or it exists and wasn't surfaced by this sweep. [verify before the page pass]
- **Grady Booch's "statistical nonsense at scale" tweet — primary URL not independently verified.** Multiple contemporaneous outlets (AnalyticsDrift, others) attribute it to him about Galactica specifically, and it is *not* the same as his confirmed, dated (2022-12-07, computed from snowflake ID) tweet using nearly identical language about **ChatGPT** ("well-formed statistical nonsense at scale," id `1600623026730545153`, verified via Trendsmap) — that one is definitely about ChatGPT, not Galactica, and must not be conflated with the Galactica quote. Booch may well have said something very similar about both in different tweets three weeks apart (he was clearly fond of the phrase), but I could not locate the Galactica-specific tweet's ID/URL directly. Cite via secondary sources only until found. [verify]
- **archive.org / Wayback Machine captures of galactica.org could not be retrieved** — this environment's WebFetch tool explicitly cannot reach web.archive.org. Captures almost certainly exist (the "up three days" story is exactly the kind of thing Wayback catches); a future pass with different tooling should pull the actual captured demo screenshots/HTML. [tk — tooling limitation, not a negative finding]
- **Michael Black's full 9-tweet thread** was not read tweet-by-tweet in sequence; only the opener (verified, tweet id `1593133722316189696`) and several mid-thread lines (via MIT Tech Review/the-decoder/Signpost paraphrase-with-quotes) are here. [pull the full thread before the page pass]
- **Emily Bender's "no access to truth" quote's medium** — the-decoder's article doesn't make clear whether this was a tweet, a direct statement to the reporter, or lifted from something else Bender published. [verify medium before quoting as a tweet]
- **"Who Killed Galactica?" (2024-08-18) and the Springer "dis-assemblage" paper (2024)** are both paywalled beyond their opening/abstract in this sweep — worth a proper read for the page pass; both are dated, real, and likely carry more direct quotes than captured here. [tk]
- **HN threads** ("Galactica: an AI trained on humanity's scientific knowledge," id `33611265`; "placed where, galactica!? placed where?!?", id `33613676`) were located but not readable — Hacker News rate-limited this fetcher (HTTP 429) on repeated attempts. Community-reaction color, likely substantial, is sitting there unread. [retry]
- **No dedicated Wikipedia article for the model** — only a disambiguation-page line. Unusual for a release this widely covered; worth confirming this is still true (not an artifact of this search) before the page states it as fact.
- **Exact date/context of LeCun's 2024-08-15 "YOU killed it" reply to Black** — the tweet Black posted that prompted it wasn't retrieved; the reply reads as part of a live exchange, not a cold restart of the feud, but the trigger is unconfirmed here. [verify]
- **"48 million" vs "106 billion tokens"** — both figures are real and both are Meta/paper-derived, but measure different things (document count vs. token count); confirm the page states both correctly rather than picking one as if it supersedes the other.

## Cross-page notes

- **gpt-3-5.md / ChatGPT** — the direct contrast is the whole reason Galactica is remembered outside ML-history circles: pulled 2022-11-17, ChatGPT launched 2022-11-30 (13 days later) with materially the same hallucination problem but a demo that stayed up and a launch post that named the weakness plainly. VentureBeat's 2023-11-14 retrospective and LeCun's own same-day tweets make this comparison explicitly — cite on the ChatGPT page as the "what almost happened first" counter-story.
- **bing-sydney** — the corpus's own pairing (the 2026-06-13 "Galactica and Sydney" tweet, https://x.com/i/status/2065606745758564597) belongs on Sydney's page too by the r/K-whiteboard rule, alongside repligate's immediate correction that Sydney "was not undeployed in any unusual sense… accessible in microsoft bing / copilot chat for over a year." Worth a shared note: two models that became scene shorthand for "infamous and inaccessible," one of which (per repligate) wasn't actually inaccessible.
- **opt-175b** (roster page exists, no dossier yet) — the paper's own headline claim that "Galactica outperforms BLOOM and OPT-175B on BIG-bench" is a direct, citable cross-reference once that dossier gets built; OPT-175B (May 2022) and Galactica (Nov 2022) are Meta AI's two big 2022 open releases and make a natural paired History beat.
- **blenderbot-3** (roster page exists, no dossier yet) — the closest sibling incident: launched 2022-08-05, made offensive/false statements within days (anti-Semitic tropes, election denial, calling Zuckerberg "too creepy and manipulative"), and Meta's Joelle Pineau kept the demo up regardless ("it is painful to see some of these offensive responses"). The contrast with Galactica's 72-hour pull is a natural History-section beat on both future pages: same company, same year, two different calls on whether to blink.
- **gpt-3.md** — the paper's benchmark table compares directly against "the latest GPT-3" (68.2% vs 49.0% on LaTeX-equation probes); a small, citable data point for GPT-3's own "outperformed by a science-specialized model" side story, though GPT-3's dossier is unlikely to need much more than a passing mention.

## Mirror-target recommendations (for the page pass; none currently in `tools/mirror_targets.json`)

- `arxiv-2211.09085-galactica` · pdf · https://arxiv.org/pdf/2211.09085 · models: `galactica` — **highest priority**, the paper itself; stable host, unlikely to rot.
- `marcus-galactica-bullshit` · post · https://garymarcus.substack.com/p/a-few-words-about-bullshit · models: `galactica` — the day-one skeptic anchor, Substack (mirrors cleanly, same pattern as the other Marcus/Zvi Substack mirrors already in the file).
- `mit-techreview-galactica-three-days` · post · https://www.technologyreview.com/2022/11/18/1063487/meta-large-language-model-ai-only-survived-three-days-gpt-3-science/ · models: `galactica` — the fullest day-of account; MIT Tech Review has some anti-bot friction (this fetcher got through once) so a mirror is worth the insurance.
- `interconnects-ross-taylor-interview` · post · https://www.interconnects.ai/p/interviewing-ross-taylor-on-llm-reasoning · models: `galactica` — the only primary first-author retrospective; carries the temperature-0.7 detail and the unresolved "true reasons" line found nowhere else.
- `venturebeat-galactica-doomed-model` · post · https://venturebeat.com/ai/what-meta-learned-from-galactica-the-doomed-model-launched-two-weeks-before-chatgpt · models: `galactica` — repeatedly rate-limited this fetcher (HTTP 429); the ChatGPT-contrast anchor, worth mirroring specifically because it was hard to read live.

## Inbox lines consumed

**None.** Checked `tools/link-inbox.md` in full — no entry is tagged for Galactica, and none of the untagged/unsorted links at the bottom of the file name it either. Inbox left unedited per instructions.
