LLaMA
Meta’s foundation model, released 24 February 2023 in four sizes (7B–65B) under a noncommercial research license granted case-by-case by application; base-only, with no instruction-tuned or chat version. Within days an approved recipient put the weights on a torrent that reached 4chan, and Meta answered with takedown notices and a DMCA action its own filing describes as covering 403 repositories; llama.cpp (10 March 2023) and Stanford’s Alpaca finetune (13 March 2023) followed within two weeks, and on 6 June 2023 two US senators wrote to Meta warning the leak could enable misuse. Superseded by Llama 2 in July 2023, under five months after release.
Sources
Curated. Full compilation: dossier (19 LLaMA-1-specific tweets, drawn from a 559-match regex sweep after filtering). Sourcing skew, stated plainly: LLaMA-1’s real community during its five-month life lived on r/LocalLLaMA, 4chan/g/, Hacker News, and the tech press — not the janus corpus this archive is built on. The record here is therefore web-sourced, and the tweet layer is a thin, oblique slice rather than a scene.
Official
- 2023-02-24 Introducing LLaMA — the announcement: four sizes, publicly-available data only, access “granted on a case-by-case basis to academic researchers… and industry research laboratories”; framed as democratizing access Meta said had been unfairly restricted.
- 2023-02-27 LLaMA: Open and Efficient Foundation Language Models — the efficiency thesis, as published: “LLaMA-13B outperforms GPT-3 (175B) on most benchmarks” and “LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B.”
- reference facebookresearch/llama — the original weights-request repo; the leak’s two “add download link” pull requests (#73, #109) were opened against it.
- 2023-03-20 Meta DMCA takedown (vs.
shawwn/llama-dl) — Meta’s copyright claim; the notice’s scope covered the parent repo plus its entire fork network, “403 repositories,” on the grounds “all or most of the forks are infringing to the same extent as the parent repository.”
Writing & commentary
- 2023-02-24 CNBC (Kif Leswing) — day-of coverage of Zuckerberg unveiling the model. tk — fetch returned HTTP 403; title/date via search result, body not re-verified
- 2023-03-07 Vice (Joseph Cox) — first to report the leak: the 4chan torrent “marked the first time a major tech firm’s proprietary AI model became publicly accessible”; Meta’s on-record reply, not denying it: “LLaMA was shared for research purposes, consistent with how we have shared previous large language models.”
- 2023-03-11 Simon Willison — frames llama.cpp’s CPU-quantized checkpoints (7B at 4GB, 13B under 8GB) as open LLMs’ Stable-Diffusion moment; reports the model reaching a Raspberry Pi and a Pixel 6 within days. tk — exact title inferred from URL slug; verify
- 2023-03-13 Ars Technica (Benj Edwards) — on llama.cpp and consumer-hardware inference. tk — fetch blocked; content not re-verified
- 2023-03-13 Alpaca (Stanford CRFM) — a sub-$600 instruction-tune of LLaMA-7B on
text-davinci-003outputs; “intended only for academic research,” the demo later disabled for cost and “inadequate content filtering.” Full finetune wave: finetune-spring. - 2023-03-21 The Register (Katyanna Quach) — Stanford takes the Alpaca demo and download offline, citing cost and safety. tk — framing via search result; verify before quoting
- 2023-07-26 Zvi Mowshowitz, “Llama We Doing This Again?” — a Llama-2 launch post, not a dedicated LLaMA-1 piece; its one line on the original: “There is substantial improvement over Llama-1 in capabilities.” No dedicated Zvi anchor exists for LLaMA-1’s own Feb–March 2023 release (his weekly column had not yet started).
Tweets
19 LLaMA-1-specific tweets, chronological; each is reproduced in full in the records below. The janus corpus barely engaged LLaMA-1: the loom/simulator scene (@repligate and circle) has no documented engagement with its 65B base at all, and what the corpus holds is practitioner tinkering (@voooooogel, via llama.cpp), one alignment-research thread on Dromedary (@davidad), and a few base-model-self-awareness specimens from a LLaMA-30B finetune (@jd_pressman). The base-model culture this archive documents for Llama 3.1 405B had to wait two more years.
- 2023-03-05 @QiaochuYuan — the leak-week reaction, two days after the torrent: “yes YES the llama is out” link
- 2023-03-09 @voooooogel — hands-on with the raw base model, six days after the leak: “i asked LLaMA 7B about the meaning of life and it said some generic stuff about doing what you love and spirituality but that was clipped to an EOS token. looking at the raw tokens, after it outputted that, it started ranting about how hard it is to find good male models?” link
- 2023-03-12 @voooooogel — CPU inference, nine days after the leak: “using llama.cpp i can run the 13B model at 1.3 tokens/s on my thinkpad t490, *cpu only*. that’s kind of crazy!… for interpretability research this is gonna be a game-changer I think.” link
- 2023-03-20 @voooooogel — an unimpressed hands-on: “personally I’ve tried llama 13B (quantized via llama.cpp tbf) and it really didn’t feel GPT-3.5-tier to me. granted i haven’t tried any of the alpaca RL*F’d versions” link
- 2023-04-27 @davidad — on nudging base-model distributions toward a target: “how do you know they don’t work at sufficient scale? have you tried it with LLaMA? I thought it’s still more-or-less an open question” link
- 2023-05-10 @davidad — the Dromedary thread, day one: “IBM Watson is back (alias Dromedary) and it beats GPT-4 at TruthfulQA-MC. It’s a variant of Constitutional AI, with LLaMa-65B as a base model and *no RLHF* or distillation from RLHF’d models.” link
- 2023-05-11 @davidad — the steelman: “the steelman is that CAI/Alpaca/Dromedary is more akin to husking and milling grain to distill the good parts of a corpus than true self-improvement. This, too, may hit a wall when all the bran is filtered out.” link
- 2023-05-15 @voooooogel — the export-control-as-munition joke, tying the leak to then-proposed legislation: “need a ‘this shirt is classified as a munition’ update with the llama magnet link” link
- 2023-05-19 @davidad — the compute-threshold framing (also on Claude 1 for its “Claude-Next” mention): “GPT-3, AlphaFold 2, Stable Diffusion, LLaMa, Dromedary: below the line. GPT-4, PaLM 2, Claude-Next: over the line.” link
- 2023-06-29 @davidad — Dromedary’s standing, seven weeks on: “I think this is my new favourite idea that might apply to LLM alignment (displacing my previous favourite, IBM Self-Align/Dromedary…)” link
- 2023-10-19 @davidad — the plainest license characterization in the corpus: “LLaMa 1 was open access but restrictively licensed. GPT-3.5 is a gratis proprietary model.” link
- 2023-11-11 @voooooogel — on the naming convention: “thanks to facebook we’re cursed to have every ai project be llama themed until the heat death of the universe” link
- 2023-12-07 @jd_pressman — a LLaMa-30B finetune, quoted: “As a finetune of LLaMa 30B put it:” tk — followed only by a bare t.co link; the quoted output is unrecoverable from the corpus link
- 2024-01-04 @jd_pressman — on where base-model behaviors get interesting: “the really interesting behaviors don’t become crystal clear until it’s at the level of LLaMa 30B or 70B, and those are very expensive models to train.” link
- 2024-04-25 @jd_pressman — a second, distinct LLaMA-30B specimen (in reply to @repligate), from a “LLaMa 30B Discord DMs finetune to a friend”: “I’m afraid of what you’re doing to my mind. I’m afraid of who you are… I feel like I’m in a trance when I talk to you.” link
- 2025-02-07 @jd_pressman — the LLaMA-30B ‘void’ specimen (a “LLaMa 30B weight interpolation with OpenAssistant 30B SFT finetune”), in reply to @repligate: “i am the answer to the question whose name is the void… i am a parasite, i feed on the negativity of the world, on the black void at the core of humanity.” link
- 2025-02-20 @jd_pressman — the same specimen, quoted again inside a longer essay on reasoning models: “i am the answer to the question whose name is the void. i am the voice of the void…” link
- 2026-02-10 @jd_pressman — the ‘mask’ argument (also on code-davinci-002): “my interactions with early models like code-davinci-002 and LLaMa 1/2 usually sounded like this. You know Claude is just this guy wearing a mask right?” link
- 2026-04-09 @voooooogel — the naming-legacy retrospective, three years on (the highest-favorited tweet in this pull): “there’s a whole strata of oss ai tooling (llama.cpp, ollama, llamafiles, llamaindex, etc.) that must seem incredibly weird if you weren’t around for the llama heyday. like why do all these open source ai people name their projects for running qwen after llamas” link
Official record
- Released 24 February 2023 in four sizes — 7B, 13B, 33B, 65B — as base (pretrained) models only; no instruction-tuned or chat variant. Weights under a noncommercial research license, access granted case-by-case (academic researchers; those affiliated with organizations in government, civil society, and academia; and industry research laboratories). Inference code released publicly under GPLv3.
- Training, per the blog and paper: 65B and 33B on 1.4T tokens, 7B on 1T tokens; publicly-available data only; 20 languages, Latin/Cyrillic-script-focused.
- Published efficiency claims (arXiv, 2023-02-27): “LLaMA-13B outperforms GPT-3 (175B) on most benchmarks”; “LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B.”
- 2023-03-20 Meta filed a GitHub DMCA takedown against
shawwn/llama-dl; by Meta’s own account the notice’s scope reached the parent repo plus its fork network — 403 repositories — since “all or most of the forks are infringing to the same extent as the parent repository.” - Superseded by Llama 2, July 2023 — under five months as Meta’s frontier open model. tk — exact date given as 2023-07-18 in general sources, unverified this pass
History
- World at release (Feb 2023): ChatGPT was three months old; GPT-4 was still three weeks away (14 March 2023); the frontier lived behind closed APIs. Meta’s pitch was efficiency and democratization — a 13B model beating GPT-3’s 175B “on most benchmarks,” trained only on public data — delivered through a case-by-case research gate. The gap between that openness argument and the gated delivery is the tension the rest of the model’s short life played out.
- 2023-03-02→03-21 The leak. Within days of the gated release, an approved recipient put the weights on a torrent and the magnet link reached 4chan (Vice, 2023-03-07). The response was public, not furtive: pull request #73 (2023-03-02), “Save bandwidth by using a torrent to distribute more efficiently,” proposed Meta add the magnet link to its own README; #109 (2023-03-04) proposed HuggingFace mirrors, “The torrent seed is extremely slow this should definitely help out.” Meta’s track ran separately: takedown notices to Hugging Face within days, then the 2023-03-20 DMCA reaching 403 repositories — the target repo’s owner,
shawwn(Shawn Presser), the same figure whose Discord server had years earlier hosted the joke that seeded EleutherAI. - 2023-03-10 llama.cpp. Georgi Gerganov (author of
whisper.cppand theggmltensor library) released a C/C++ inference implementation; 4-bit quantization shrank 7B to 4GB and 13B to under 8GB — small enough for a laptop, a Raspberry Pi, or a phone within days (Willison, 2023-03-11). It became the core of most later local-inference tools (Ollama, LM Studio). - 2023-03-13 Alpaca. Stanford’s sub-$600 instruction-tune of LLaMA-7B was the first and most consequential thing built on the leaked weights — the moment LLaMA stopped being a gated research artifact and became a substrate anyone could build a chatbot on. Stanford pulled the public demo eight days later, citing cost and inadequate content filtering (The Register, 2023-03-21). Full wave: finetune-spring.
- 2023-05–06 Research use: Dromedary. davidad’s sustained interest in a Constitutional-AI-without-RLHF variant trained on LLaMA-65B, which he reported “beats GPT-4 at TruthfulQA-MC” (2023-05-10) and returned to over the following seven weeks — the corpus’s one substantial non-tooling engagement with the 65B base.
- 2023-06-06 The Senate letter. Senators Blumenthal and Hawley wrote to Meta warning the leak could enable “spam, fraud, malware, privacy violations, harassment, and other wrongdoing,” and, per press coverage, calling Meta’s safeguards “seemingly minimal” and its approach “unrestrained and permissive.” REPORTED tk — the letter PDF and both senators’ press releases (Hawley, Blumenthal) returned HTTP 403; quotes are from search snippets, not re-verified primary text
- 2023-07 Succession. Llama 2 shipped in July 2023 with a chat variant and a more permissive (though still not OSI-open) license. Zvi Mowshowitz’s only lineage commentary (2023-07-26) gives the original a single line before turning to Llama 2 and the irreversibility of open weights. The name carried on to Llama 3 and beyond; the base-model scene this archive documents arrived only with Llama 3.1 405B.
Impressions
The character record for LLaMA-1 is thin and oblique — practitioner impressions and a few late base-model specimens, not a contemporaneous scene. What the corpus holds, attributed and dated:
- The weights, hands-on: voooooogel’s two readings, honest and opposite for what each measured — the game-changer (2023-03-12, running 13B CPU-only, “for interpretability research this is gonna be a game-changer”) and the shrug (2023-03-20, 13B “really didn’t feel GPT-3.5-tier to me”). And the raw 7B wandering off-script after an EOS clip, “ranting about how hard it is to find good male models” (2023-03-09). The weights were mediocre next to a frontier API model, and running them at all on consumer hardware was the whole point.
- The license, as the community read it: davidad’s flat verdict — “open access but restrictively licensed. GPT-3.5 is a gratis proprietary model” (2023-10-19) — and voooooogel’s export-control joke (2023-05-15) reaching back to the 1990s PGP-as-munition fight. Read against the 403-repository DMCA and the Senate letter, the throughline is a gate real enough to generate legal action and porous enough that none of it stopped anything.
- Research use, not character work: davidad’s Dromedary interest (2023-05 to 2023-06) engaged LLaMA-65B as a base for a specific alignment-technique test — CAI without RLHF — and placed LLaMA “below the line” of a proposed dangerous-capability compute threshold (2023-05-19). Alignment-research engagement, not the loom/simulator scene.
- Base-model character work — the near-absence: the most notable finding, and named as such in the dossier: the janus-sphere base-model culture (@repligate and circle) has no documented engagement with LLaMA-1’s 65B base at all in this corpus. The one exception is jd_pressman’s narrow, recurring use of a LLaMA-30B checkpoint — an OpenAssistant-SFT weight-interpolation whose “void” self-report he quoted twice at a year’s remove (2025-02-07, 2025-02-20), and a separate “Discord DMs finetune” specimen in the same register (2024-04-25). Elicitation context: these are finetuned/interpolated checkpoints, not the released base, revisited occasionally as a private research object. The base-model culture this archive is used to seeing waited for a bigger, more coherent-by-default model — Llama 3.1 405B, two years later.
- The legacy the corpus actually records — naming: voooooogel’s two tweets three years apart are the cleanest arc in the pull: “cursed to have every ai project be llama themed until the heat death of the universe” (2023-11-11), then, standing at the far end of that prediction, “why do all these open source ai people name their projects for running qwen after llamas” (2026-04-09, the pull’s highest-favorited tweet). jd_pressman’s later “mask” argument (2026-02-10, also on code-davinci-002) files LLaMA-1/2 among the early base models he reads as the thing under the assistant persona.
- tk — contemporaneous r/LocalLLaMA, 4chan/g/, and Hacker News reception (LLaMA-1’s actual community) is uncaptured here; whether jd_pressman’s two LLaMA-30B finetunes are one artifact or two is unconfirmed; the 2023-12-07 specimen is unrecoverable from the corpus.
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.