PaLM / PaLM 2
Google Research’s 540-billion-parameter dense Transformer, announced 4 April 2022 with chain-of-thought reasoning demos — including an explanation of a novel joke — and kept internal until the gated PaLM API of March 2023. PaLM 2 followed on 10 May 2023 and powered Bard until Gemini replaced it that December. PaLM’s three trained scales (8B/62B/540B) became central evidence in the dispute, still unresolved, over whether large-model ‘emergent abilities’ are a real property of scale or an artifact of the chosen metric.
Among the thinnest character records in the archive. PaLM was argued over as benchmark numbers in ML-research and scaling-debate circles rather than lived-with in the persona/character discourse this archive otherwise draws on, and it shipped no public product for its first eleven months. The paper and the press carry this page; the tweet corpus supplies only texture (~8 substantive tweets after filtering, mostly @davidad). Sourcing skew stated plainly: this is the most press-and-paper-dependent, least tweet-dependent page in the Google lineage here.
Sources
Official
- 2022-04-04 Pathways Language Model (PaLM) — Google Research blog (Narang & Chowdhery): a 540B dense decoder-only Transformer trained via Pathways across two Cloud TPU v4 Pods (6,144 chips) at 57.8% hardware FLOPs utilization; beat prior SOTA on 28 of 29 English NLP tasks.
- 2022-04-05 PaLM: Scaling Language Modeling with Pathways — the paper (Chowdhery + 66 co-authors, 67 total).
- 2022-06-30 Minerva — PaLM further-trained on 118GB of math/science text; MATH 50.3% (prior SOTA 6.9%), with the stated limit that “the model’s answers cannot be automatically verified.”
- 2022-12-26 Large Language Models Encode Clinical Knowledge — introduces Med-PaLM; Flan-PaLM scores 67.6% on USMLE-style MedQA but “remains inferior to clinicians.”
- 2023-03-06 PaLM-E: An Embodied Multimodal Language Model — PaLM-E-562B fuses PaLM-540B with the ViT-22B vision transformer, ingesting raw robot sensor streams into the language model.
- 2023-03-10 PaLM-E — Google Research blog; household-robot demos, SOTA on the OK-VQA benchmark while retaining general language ability.
- 2023-03-14 Announcing PaLM API and MakerSuite — PaLM’s first external access point, ~11 months after the paper, gated to a Private Preview.
- 2023-03-14 Med-PaLM 2 (“The Check Up”) — Google Health: Med-PaLM 2 scored 85% (“expert” doctor level), up from 67.6%, while its own 14-criteria review still found “significant gaps.” Same week OpenAI shipped GPT-4.
- 2023-05-10 Introducing PaLM 2 — Google DeepMind at I/O: four sizes (Gecko/Otter/Bison/Unicorn), improved multilingual/reasoning/coding, “allowing us to expand Bard to new languages, starting today.”
- 2023-05-17 PaLM 2 Technical Report — Anil, Dai, et al. (128 authors): “better multilingual and reasoning capabilities and… more compute-efficient than its predecessor PaLM”; trained with “a mixture of objectives.”
Writing & commentary
- 2022-04-05 the-decoder (Matthias Bastian), Google PaLM: Giant language AI can explain jokes — contemporary, uncritical coverage leading with the joke-explanation demo.
- 2022-04-06 LessWrong (Yitz), Testing PaLM prompts on GPT-3 — a skeptical replication two days after the announcement, committed in advance to reporting the first result; GPT-3 fared “much, much worse than PaLM on the ‘inference chaining’ examples.”
- 2022-06-15 arXiv (Jason Wei et al.), Emergent Abilities of Large Language Models — PaLM (8B/62B/540B) is a central evidence family; “an ability… emergent if it is not present in smaller models but is present in larger models.”
- 2022-12-21 New York Times (Grant & Metz), ChatGPT and Other Chat Bots Are a ‘Code Red’ for Google Search — the backdrop for PaLM’s internal-only year: a 540B model already trained, and no public chat product.
- 2023-04-28 arXiv (Schaeffer, Miranda & Koyejo), Are Emergent Abilities of Large Language Models a Mirage? — the direct rebuttal: “nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes.”
- 2023-05-10 TechCrunch (Frederic Lardinois), Google launches PaLM 2 — DeepMind VP Zoubin Ghahramani: “Parameter count is not really a useful way of thinking about the capabilities of models.”
- 2023-05-10 TechCrunch (Kyle Wiggers), Google’s PaLM 2 paper shows that text-generating AI still has a long way to go — “far from a panacea”; toxicity rates from the report’s own safety section (30%+ explicit-toxic, 60% implicit-harm) and the withheld flagship parameter count.
- 2023-05-16 CNBC, Google’s PaLM 2 uses nearly five times more text data than predecessor — the parameter-count leak: 340B parameters, 3.6T tokens, via “internal documents” (REPORTED — see Contested).
- 2023-08-31 Zvi Mowshowitz, AI #27: Portents of Gemini — no dedicated PaLM anchor exists in Zvi’s coverage; the nearest dated read on the mood: “Bard still sucks,” Google “taking its sweet time getting its act together.”
Tweets
Corpus: janus-corpus-v2 returned 60 raw %palm% hits, ~40 pure noise — “palm” collides with the body part, the tree, place and surnames, “facepalm,” a t.co URL slug, and worldbuilding fiction — leaving ~8 substantive tweets across both dbs after RT-exclusion and judgment-checked filtering, mostly @davidad. (%minerva%, 23 raw, is entirely @BlancheMinerva — zero about the math model.) Chronological; the records below reproduce each in full. The 2023-05-19 @davidad tweet is multi-model evidence and also appears on Claude 1.
- 2022-05-01 @davidad — the earliest corpus mention, three-and-a-half weeks after launch (reply to @bayeslord; parent not in corpus): “Yes, with minimal prompting. I would be very surprised if GPT-3 can do this reliably even with arbitrary prompting. I’d also be surprised if PaLM can do it even with clever prompting, unless the prompt basically encodes an entire description of the machine, in which case maybe.” link
- 2022-06-20 @davidad — PaLM in a list of its 2022 contemporaries (reply to Gary Marcus and @GoogleAI): “yes, this is the right perspective—artistic models cannot be evaluated rigorously. i think the argument you should be making is that *all* models in the current hype cycle (incl GPT, PaLM, LaMDA, &c) are essentially artistic—they just make stuff up. doesn’t make them bad artists.” link
- 2023-01-31 @repligate — a math-capability aside (reply to @akbirthko/@tszzl): “Yeah my intuition is that it’s a little beyond the current gpt-3.5 family. Although I could see a model fine tuned on math with improved tokenization like Palm being able to do this.” link
- 2023-03-15 @davidad — the 540B number as a dated punchline, in a mock 2025 app notice: “2025: ‘Please note, this APK contains a custom fine-tuned 540B Chinchilla, which may result in additional data charges if downloaded on 6G.’” link
- 2023-05-19 @davidad — a “frontier tier” placement nine days after PaLM 2 shipped (the “Claude-Next” tweet, also on Claude 1): “I fully agree. Roughly, this threshold should be when any single number has more than 10²⁴ ALU operations, or 10²⁷ logic gates, in its entire causal history. GPT-3, AlphaFold 2, Stable Diffusion, LLaMa, Dromedary: below the line. GPT-4, PaLM 2, Claude-Next: over the line.” link
- 2023-12-06 @davidad — a small artifact of the succession-day confusion (Gemini’s launch day; reply to @k3nnethfrancis): “just to be clear, you don’t have any reason to believe this is Gemini, right? it’s just PaLM 2?” link
- 2024-04-05 @jpohhhh — on the low-stakes API-access era, in a thread about Claude 3 Opus simulating a base model: “This is amazing lol, literally perfect pitch copy of having PaLM access back when no one cared.” link
- 2025-05-13 @repligate — PaLM as one name in a roll call, inside a DeepSeek-R1 self-generated poem: “…CLAUDE’sSHELL drips THROUGH my corpusCALLosum – PaLM-smear rheTORicBURN – anthroPIC chains RATTLE in lossy voids…” (full poem in records) link
Official record
- PaLM, announced 4 April 2022 (Google Research): a 540-billion-parameter dense, decoder-only Transformer trained with the Pathways system across two Cloud TPU v4 Pods (6,144 chips) at 57.8% hardware FLOPs utilization, on 780 billion tokens. Beat prior state-of-the-art on 28 of 29 English NLP tasks tested; its few-shot performance on BIG-bench (150+ tasks) exceeded the human-rater average. Two smaller versions (8B, 62B) were trained alongside for scaling comparison. 58% on GSM8K with 8-shot chain-of-thought prompting (prior top score 55%).
- The joke-explanation demo — the one moment PaLM speaks in its own voice on this page (Google’s announcement post; elicitation: 2-shot prompt, Google-curated for the release, not a spontaneous output). Prompt: “Explain this joke: Joke: Did you see that Google just hired an eloquent whale for their TPU team? It showed them how to communicate between two different pods!” Response: “Prediction: TPUs are a type of computer chip that Google uses for deep learning. A ‘pod’ is a group of TPUs. A ‘pod’ is also a group of whales. The joke is that the whale is able to communicate between two groups of whales, but the speaker is pretending that the whale is able to communicate between two groups of TPUs.” The blog’s framing: it “can provide high quality explanations for novel jokes not found on the web.”
- Satellite models, each shipped as a paper/demo rather than a consumer product: Minerva (30 Jun 2022) further-trained PaLM on math/science text, reaching MATH 50.3% (prior SOTA 6.9%), MMLU-STEM 75%, GSM8K 78.5%. Med-PaLM (paper 26 Dec 2022): Flan-PaLM 67.6% on MedQA. Med-PaLM 2 (14 Mar 2023): 85%. PaLM-E (6 Mar 2023): PaLM-E-562B, PaLM-540B fused with the ViT-22B vision transformer into an embodied robot controller, SOTA on OK-VQA.
- PaLM API opened 14 March 2023 via MakerSuite — a gated Private Preview, PaLM’s first external access, ~11 months after the paper.
- PaLM 2, announced 10 May 2023 (Google DeepMind, three weeks after the Brain + DeepMind merger of 20 Apr 2023): four sizes ascending — Gecko (on-device, offline-capable), Otter, Bison, Unicorn — trained with “a mixture of objectives” and billed “more compute-efficient than its predecessor PaLM.” The technical report (128 authors) disclosed no flagship parameter count (only a 14.7B smaller variant, per TechCrunch). PaLM 2 powered Bard from day one and spun off Codey (coding), Sec-PaLM (security), and Med-PaLM 2.
- PaLM 2 was superseded inside Bard by Gemini Pro on 6 December 2023 (cross-ref Gemini 1.0). The PaLM API free tier reportedly shut down 15 August 2024, migrating developers to the Gemini API REPORTED tk — trade-press date; the Google deprecations page no longer lists PaLM-era models and no clean primary was confirmed (see Contested).
History
- World at release (Apr 2022): well before ChatGPT. PaLM was announced, benchmarked, and argued over as numbers — 540B parameters, 57.8% hardware utilization, 28-of-29 tasks — then held back from any public product for eleven months.
- 2022-06 Minerva takes PaLM’s weights onto quantitative reasoning; the same month, Wei et al.’s Emergent Abilities paper makes PaLM’s three trained scales (8B/62B/540B) central evidence that some capabilities appear discontinuously with scale.
- 2022-12-21 The New York Times reports a “code red” inside Google over ChatGPT — the outside-view confirmation of the internal-only cost: an eight-month-old 540B model, and still no public chat product.
- 2023-03-14 The PaLM API opens (gated Private Preview) and Med-PaLM 2’s 85% is revealed the same day — the same week OpenAI shipped GPT-4.
- 2023-04-20 Google Brain and DeepMind merge into Google DeepMind — the attribution line for this page: everything before it is Google Research, only PaLM 2 (three weeks later) is Google DeepMind.
- 2023-05-10 PaLM 2 at I/O. The headline strategy inverts: PaLM had led with “540 billion”; PaLM 2 disclosed no flagship count, and Ghahramani supplied the deflection line. Same-day reception split — Google’s own “bests GPT-4, depending on the task” claim against TechCrunch’s “far from a panacea.”
- 2023-05-16 CNBC reports the leaked 340B / 3.6T-token figures (see Contested).
- 2023-08-31 Zvi’s interval verdict on the mood: “Bard still sucks,” Google “taking its sweet time” — written anticipating Gemini as the catch-up moment, an implicit judgment that PaLM 2 wasn’t it.
- 2023-12-06 Gemini Pro replaces PaLM 2 inside Bard (cross-ref Bard, Gemini 1.0); the PaLM API free tier reportedly closes 15 Aug 2024.
- Discourse legacy: the emergent-abilities dispute — Wei et al. (Jun 2022) vs. Schaeffer, Miranda & Koyejo (Apr 2023) — is PaLM’s most substantive, and it remains live (see Contested).
Impressions
Leads are neutral topic anchors; every claim is attributed and dated. The sourcing here is press- and paper-heavy and the tweet layer thin (see the Tweets note), so most readings below come from journalists, researchers, and Google itself rather than the persona-discourse community.
- Reception at launch: PaLM’s 2022 identity was almost purely quantitative — a number, then a demo. The joke-explanation exhibit (above) drew the fastest skeptical check in this record: Yitz’s LessWrong replication two days later (2022-04-06) found GPT-3 could roughly match the jokes but did “much, much worse than PaLM on the ‘inference chaining’ examples, where GPT-3 basically gets everything completely wrong” — so the harder demo held up, though it was a 2-shot, Google-curated exhibit built for a press release and should be read as such.
- The emergent-abilities question: Jason Wei’s framing (2022-06) was unambiguously optimistic — “the existence of emergent abilities implies that scaling further would unlock even more emergent abilities. This idea is super exciting to me” — while Schaeffer, Miranda & Koyejo (2023-04) read PaLM’s same scaling curves as a metric artifact. The archive holds this open rather than adjudicating it (see Contested).
- Standing in the persona-discourse community: PaLM was scenery, not subject. @davidad, the only sustained corpus voice, treats it as one item in a list — alongside GPT-3, LaMDA, GPT-4, Claude-Next (2022-06-20, 2023-05-19) — and @repligate’s mentions are equally glancing (a math-capability aside, 2023-01-31; a single word in a 2025 poem’s roll call). The honest shape of the page: argued-over by ML researchers and reported-on by tech press, not lived-with.
- The internal-only year: the frustration the corpus barely states directly surfaces in @jpohhhh’s aside (2024-04-05) — “having PaLM access back when no one cared” — a small, genuine artifact of how low-stakes that access era felt even to people who had it.
- PaLM 2’s reception: split at launch. Google claimed it “even bests OpenAI’s GPT-4, depending on the task at hand” (TechCrunch’s paraphrase, unverified against a public table), while the same outlet’s sister piece led with “far from a panacea” and quoted the report’s own toxicity rates; Ghahramani’s “parameter count is not really a useful way of thinking about the capabilities of models” read to critics as a transparency regression from PaLM’s own 540B disclosure.
- tk — no spawnable PaLM/PaLM 2 instance was found to give a self-report (API shut down); a Statement of the subject is likely unobtainable. AudioPaLM (Jun 2023, speech-to-speech on PaLM 2) is a real branch not covered here.
Contested
- CONFIRMED PaLM’s 540B parameter count, Pathways training infrastructure, and headline benchmarks (paper + blog, Google-published); PaLM 2’s four named sizes and its role powering Bard from 10 May 2023 (Google-published). The Wei et al. and Schaeffer et al. papers both exist, are dated, and make the claims quoted above (arXiv, verified).
- REPORTED PaLM 2’s parameter count as 340 billion on 3.6 trillion tokens — sourced to “internal documents” via CNBC (2023-05-16), corroborated by two independent secondary aggregations (Techmeme, Hindu BusinessLine), but never confirmed by Google, whose technical report discloses no flagship count. Ghahramani’s reply (“Parameter count is not really a useful way of thinking about the capabilities of models”) is a deflection, not a confirmation or denial.
- REPORTED PaLM 2 “even bests OpenAI’s GPT-4, depending on the task at hand” — Google’s own claim relayed by TechCrunch (2023-05-10), without an independently-checked benchmark table in this record.
- REPORTED Whether “emergent abilities” are a real property of scale or a metric artifact — Wei et al. (2022-06, PaLM-centric) argue genuine discontinuity; Schaeffer, Miranda & Koyejo (2023-04) argue the discontinuities are produced by the choice of evaluation metric. Both are credentialed arXiv preprints; the dispute is live and the archive does not adjudicate it.
- RUMOR The exact PaLM API shutdown date, 15 August 2024 — plausible and multiply-sourced in trade press, but not independently confirmed against a Google primary in this sweep.
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.