Mixtral 8x7B
Released Dec 2023 as a bare magnet link — no blog post, no benchmarks — containing a frontier-adjacent sparse MoE. The release format itself became a genre.
Scope: this page covers the Mixtral sparse-MoE release line — Mixtral 8x7B (magnet 8 Dec 2023, announced 11 Dec 2023) and Mixtral 8x22B (magnet 10 Apr 2024, announced 17 Apr 2024) — as one continuous story: identical magnet-link-then-blog-post choreography, the same Apache 2.0 open-weights posture, and a corpus that repeatedly discusses the two checkpoints together. 8x22B got no dedicated paper (blog post and model card only), a real difference from 8x7B’s treatment. Checkpoint-specific facts are marked as such.
Sources
Official
- 2023-12-08 Magnet-link release (8x7B) — @MistralAI’s entire announcement is a raw BitTorrent magnet URI, no text and no benchmarks:
magnet:?xt=urn:btih:5546272da9065eddeb6fcd7ffddeef5b75be79a7&dn=mixtral-8x7b-32kseqlen… - 2023-12-11 Mixtral of Experts — the blog post three days later: 46.7B total / 12.9B active parameters, 8 experts with top-2-per-token routing, 32k context, Apache 2.0; claims it outperforms Llama 2 70B with 6× faster inference and matches or beats GPT-3.5, the Instruct variant billed as “the best open-source model.”
- 2024-01-08 Mixtral of Experts (paper, 25 authors) — confirms 47B total / 13B active parameters; states the Instruct variant “surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro, and Llama 2 70B – chat model on human benchmarks.”
- 2024-04-10 Magnet-link release (8x22B) — the same choreography repeated: a bare magnet URI,
magnet:?xt=urn:btih:9238b09245d0d8cd915be09927769d5f7584c1c9&dn=mixtral-8x22b… - 2024-04-17 Cheaper, Better, Faster, Stronger — the 8x22B blog post one week after the magnet: 141B total / 39B active parameters, 64k context, Apache 2.0, native function-calling. “We believe in the power of openness and broad distribution to promote innovation and collaboration in AI.” No dedicated arXiv paper was published for 8x22B.
Writing & commentary
- 2023-06-20 The GPT-4-is-a-MoE rumor — George Hotz’s and Soumith Chintala’s claim that GPT-4 is an ~8×220B mixture-of-experts circulates on HN/Twitter; unconfirmed by OpenAI then or since. Background: the rumor Mixtral would later give the public something concrete to point at.
- 2023-12-08 @ericjang11, same-day: “mistral’s brand is already becoming one of my favorites… releases 87GB torrent containing 8x 7B MoE model via tweet, refuses to elaborate” — tweet (not in the local corpus). And @jeremyphoward, deadpan: “nbd just @MistralAI dropping a magnet link or something” — tweet.
- 2023-12-11 VentureBeat and Voicebot.ai — day-of trade press centered on the release-format choice; Voicebot explicitly frames the $415M raise and the free-torrent drop as one story.
- 2023-12-11 Together AI — “Can you feel the MoE?” — launch-day serving partner, “over 100 tokens per second… the fastest performance at the lowest price.”
- 2023-12 Fireworks AI — the title itself is evidence of the magnet-link culture: an inference provider had it running “even before the official release.” tk — exact publication date
- 2023-12-14 Eric Hartford — dolphin-2.5-mixtral-8x7b — the uncensored finetune, six days after the base weights; by Hartford’s own account a system prompt offering a fictional “$2000 tip” for compliance, built to be “over-the-top uncensored,” with a warning to downstream operators to add their own alignment layer.
- 2023-12-18 Simon Willison — running Mistral models in your terminal — the fullest day-of technical reaction: names the magnet drop, calls Mixtral “the first truly convincing openly licensed implementation of this architecture I’ve seen,” and worries the hosting “race to the bottom… actively disincentivizes future open model releases from Mistral.”
- 2023-12-21 Zvi Mowshowitz — AI #43 — sole launch-week mention: “Mixtral offering their API for free while supplies last.”
- 2024-01-18 Zvi Mowshowitz — AI #47 — on the “Phixtral” hack, quoting Shital Shah: “you can just slap in pre-trained models as ‘experts’ in Mixtral.” (in-text numbered #47 but the URL slug reads “ai-48” — reproduced as found)
- 2024-01 Nous-Hermes-2-Mixtral-8x7B-DPO — Mixtral’s fastest-adopted finetune substrate after Dolphin; full context on Nous-Hermes. tk — exact January date
- 2024-02 Hacker News — Groq runs Mixtral 8x7B-32k — the ~430–500 tokens/s throughput demo that put Groq’s custom silicon on the wider field’s map, months before its later fame. tk — exact demo date; primary Groq post if one exists
- 2024-04-10 Simon Willison — Mistral tweet a magnet link for mixtral-8x22b — names the pattern four months on: “their now standard operating procedure of tweeting out a raw torrent link”; this one “a whole lot bigger (a 281GB download).”
- 2026-06-17 Sifted — Mistral CEO pitches open source AI — the “European champion”/sovereignty frame in its most explicit form; Arthur Mensch positions Mistral as existing “outside of centralised control exercised by states or corporations.” Cited as the frame’s throughline, not day-of Mixtral coverage.
Tweets
18 corpus tweets after RT-filter and cross-db dedup (11 primary + 7 supplement); the curated selection below is chronological, and the records below will reproduce the cited tweets in full once the dossier is wired. Sourcing skew, stated plainly: Mixtral’s real launch-era community was Hacker News, r/LocalLLaMA, the inference-hosting trade press, and the Hugging Face finetune ecosystem — not this corpus’s Twitter/X scene. What the corpus captures is a narrow slice: two interpretability-minded accounts (@repligate, @voooooogel) using Mixtral base as one test subject among several, plus @jd_pressman using it as agentic-tooling substrate. The release-format, MoE-rumor, finetune, and inference-speed stories are web-sourced above, not corpus-sourced. Plain x.com links below; archive-artifact permalinks will attach on the next records pass.
- 2023-12-25 @jd_pressman — early hands-on, Christmas Day, 17 days after the magnet: “Mixtral has noticeably different biases to LLaMa 2 70B… It can’t write it, but it can write an exegesis of it.” link
- 2023-12-31 @voooooogel — the torrent as mob-boss heist, three weeks after the drop (♥16): “…that’s the torrent for the Mixtral weights! they’re handin’ ’em out like candy!… researchers these days… no respect for the game i tell you…” (full text in records) link
- 2024-01-16 @solarapparition — scaling skepticism, mid-thread: “best open model, Mixtral, uses MoE, which is not new and what GPT-4 model probably uses.” link
- 2024-01-23 @davidad — the risk-framing pole (note the ‘Mixtral-8.6T’ figure matches no shipped model — reads as hypothetical extrapolation): “An open-source general-purpose Llama7 or Mixtral-8.6T could potentially be scaffolded like ChaosGPT by 1 angry teen.” link
- 2024-01-29 @voooooogel — poking the MoE mechanism, on why repeng control vectors won’t transfer cleanly: “since the control vectors are pulled off the last token, they’re only based on the 2 experts used for that token, so when applied to other experts they don’t work right. probably need a vector per expert” link
- 2024-04-25 @repligate — the interpretability-specimen thread, base-model cosmology: “base models trained on recent data like mixtral, which exhibit more overt latent situational awareness…” link
- 2024-07-21 @jd_pressman — word-game struggles (‘Mixtral-large’ is ambiguous — possibly Mistral Large, a typo, or a larger Mixtral checkpoint; unresolved): “I noticed that Mixtral-large really struggled to play this Binglish word game unless I had exactly the right prompt.” link
- 2024-09-06 @jd_pressman — the pull’s highest-favorited tweet (♥56), 8x22B as agentic substrate: “Optimizing Weave-Agent for LLaMa 3.1 405B and (later) Mixtral 8x22B is the first time I think I’ve really experienced this firsthand in a deep way. You come to realize these models have deep aesthetic preferences your program will conform to if you want understanding from it.” link
- 2024-09-13 @repligate — still a comparison specimen a year on (the referent ‘it’ depends on an un-transcribed image — see gaps): “Even mixtral and 405 base do it (and I suspect every other new base model). If Mistral (instruct?) doesn’t do it, it’s an interesting anomaly.” link
- 2024-12-16 @voooooogel — a year later, still the go-to open MoE to experiment on: “and then mixtral (or 4base if you can swing it) if you want to try MoE.” link
- 2025-11-13 @voooooogel — on 8x7B’s close-reading, two years on: “like this was mixtral 8x7B but imo that points to them being smarter than you think, superhuman at this even. it’s just not adderall taskboi coder smarts” link
- 2025-12-01 @lu_sichu — canon marker; a satirical model catalog names both checkpoints in passing (the last one fictional; also cited on Claude 2.1): “Mixtral-8x7B, Mixtral-8x22B, Mixtral-8x34B-‘powered by spite’” link
- 2025-12-29 @voooooogel — the elegiac close of the fourth-wall thread, two years old and infrastructure-obsolete: “mixtral base isn’t hosted anymore afaik and i don’t have the time to get it running right now” (full text in records) link
Official record
- Mixtral 8x7B — announced by magnet link 8 Dec 2023, blog post 11 Dec 2023: a sparse mixture-of-experts with 46.7B total / 12.9B active parameters (8 experts, top-2-per-token routing), 32k context, released under Apache 2.0. Deployed to Mistral’s La Plateforme beta under the endpoint alias
mistral-small. The Instruct variant scored 8.30 on MT-Bench; Mistral billed it as “the best open-source model.” - Benchmarks as published — Mistral claimed 8x7B outperforms Llama 2 70B on most benchmarks with 6× faster inference and matches or beats GPT-3.5. The paper (8 Jan 2024) states the Instruct model “surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro, and Llama 2 70B – chat model on human benchmarks.”
- 2024-01-28 Independent replication: LMSYS Chatbot Arena ranked Mixtral-Instruct 8th overall on its human-preference Elo leaderboard, ahead of GPT-3.5-Turbo, Gemini Pro, Claude-2.1, and Llama 2 70B-chat. tk — primary lmsys.org source not located this pass; ranking cited via a secondary summary, verify before hardening
- Mixtral 8x22B — announced by magnet link 10 Apr 2024, blog post 17 Apr 2024: 141B total / 39B active parameters, 64k context, Apache 2.0, with native function-calling. No dedicated arXiv paper was published — a blog post and model card are the only primary technical sources found, a real difference from 8x7B’s treatment.
- Lifecycle — being Apache 2.0 open-weights, both checkpoints remain downloadable and there was no formal deprecation. By late 2025 the base model was, per one practitioner, no longer conveniently hosted (@voooooogel, 2025-12-29: “mixtral base isn’t hosted anymore afaik”) — infrastructure attrition rather than a retirement ceremony.
History
- The rumor it made checkable. Since 2023-06 GPT-4 had been rumored (George Hotz, echoed by Soumith Chintala) to be an ~8×220B mixture-of-experts — a claim no one outside OpenAI could verify, because GPT-4 was closed. Mixtral 8x7B arrived six months later as an actually downloadable 8-expert MoE, the field’s first hands-on specimen of the architecture everyone had been speculating about. Simon Willison, 2023-12-18: “GPT-4 has long been rumored to use a mixture of experts architecture, and Mixtral is the first truly convincing openly licensed implementation of this architecture I’ve seen.”
- The magnet link as method. The 2023-12-08 announcement was a bare magnet URI with no text; reaction was genre-aware the same day (Eric Jang: “refuses to elaborate”; Jeremy Howard, deadpan: “nbd just @MistralAI dropping a magnet link or something”). By the second occurrence (8x22B, 2024-04-10) Willison named it outright: “their now standard operating procedure of tweeting out a raw torrent link.” Two data points made a pattern; this corpus never shows anyone surprised by it again.
- Release and raise, same day. The 2023-12-11 blog post landed the day Mistral closed a $415M Series A led by a16z at a ~$2B valuation (TechCrunch: “Paris-based OpenAI rival”); Voicebot.ai framed the funding and the free-torrent drop as one story. The entanglement of open-weight release with financing and geopolitical signaling hardened over the following years rather than being argued in December 2023 (see Impressions).
- The finetunes, within a week. 2023-12-14 Eric Hartford shipped dolphin-2.5-mixtral-8x7b, an explicitly uncensored finetune; in 2024-01 Nous Research followed with Nous-Hermes-2-Mixtral-8x7B-DPO (full context: Nous-Hermes). This is the same open-weights-to-companion diaspora this archive documents for GPT-J and Llama 2 — the character-projection energy attached to these finetunes, not to Mixtral as shipped (see Impressions).
- The serving war. Inference providers raced Mistral’s own follow-up post: Fireworks advertised Mixtral “faster, cheaper, even before the official release”; Together shipped “over 100 tokens per second” on launch day; and in 2024-02 Groq’s ~430–500 tokens/s demo (per Hacker News) reframed fast open-model inference months before Groq’s later fame — Mixtral as the proving ground. Willison’s worry that this “race to the bottom” would disincentivize future open releases didn’t fully materialize (8x22B shipped four months later), but the margin-capture mechanism he named is exactly what the Together/Fireworks/Groq race demonstrated.
- The mechanism, hobbyist-hackable. Within six weeks, the “Phixtral” hack slotted pretrained Phi-2 checkpoints directly into Mixtral’s expert slots without further training (Zvi, 2024-01-18, quoting Shital Shah: “you can just slap in pre-trained models as ‘experts’ in Mixtral”) — the routing skeleton became something the scene could repurpose by hand.
- 8x22B repeats the pattern. 2024-04-10 magnet, 2024-04-17 blog — larger (141B/39B, a 281GB download) but no dedicated paper. The choreography was by then expected.
- Afterlife as research instrument. Through 2024–2025 @repligate and @voooooogel kept returning to Mixtral base as one interpretability specimen among several (alongside Llama 3.1 405B base) in an ongoing base-model situational-awareness thread, continuing as late as 2025-12 — by which point the weights were no longer conveniently hosted. By 2025-12 Mixtral was a stable enough name to anchor a two-year-old joke’s model-canon catalog without explanation; by 2026-06 Mistral’s CEO was invoking the sovereignty frame the magnet-link ethos only implied.
Impressions
Character claims only, attributed and dated. Given the sourcing skew above, this section leans on a narrow interpretability/practitioner slice of the scene and on web reception; it is not a broad-community read.
- Release reception: the format, not the model, was the event. The scene treated the magnet drop as already-canonical Mistral behavior on its first occurrence (Jang, Howard, above), and within three weeks it was jokeable — @voooooogel’s mob-boss dialogue (2023-12-31) built an entire heist bit around “the torrent for the Mixtral weights! they’re handin’ ’em out like candy!”
- As an architecture specimen: the scene poked at the MoE mechanism itself, not just outputs — @voooooogel theorizing (2024-01-29) that repeng control vectors “are pulled off the last token, they’re only based on the 2 experts used for that token, so when applied to other experts they don’t work right,” and the Phixtral hack treating expert slots as swappable. A year on it was still the reference open MoE to experiment on (@voooooogel, 2024-12-16: “then mixtral… if you want to try MoE”).
- As a base-model interpretability subject: @repligate cited it (2024-04-25) as a newer-generation base model exhibiting “more overt latent situational awareness” within his “IRL prior” vs. “imaginal prior” cosmology; @voooooogel later ran it against Llama 3.1 405B base in a hedged, sample-size-arguing thread about whether base models “recognize” fourth-wall-breaking text — and, separately, defended 8x7B’s close-reading as “superhuman at this even. it’s just not adderall taskboi coder smarts” (2025-11-13). Technical and hedged, not reverent.
- As practitioner substrate: @jd_pressman’s register is hands-on and comparative rather than evaluative — Mixtral has “noticeably different biases to LLaMa 2 70B” (2023-12-25), and optimizing Weave-Agent around 8x22B taught him “these models have deep aesthetic preferences your program will conform to if you want understanding from it” (2024-09-06, the pull’s highest-favorited tweet).
- Where the character energy went: to the finetunes, not the base. This archive documents an open-weights-to-companion diaspora for GPT-J and Llama 2; Mixtral joined it immediately (Dolphin’s uncensoring, Nous-Hermes’s DPO tuning) — but the affection and character-projection attached there, not to Mixtral 8x7B/8x22B as base or instruct releases.
- The absence, checked deliberately (not merely unqueried): the dossier reports finding no companion-model or loomed-persona reading of Mixtral itself — no “look what it said” screenshots, no character voice. Its corpus presence is practitioner tooling and interpretability specimen, full stop — a structural parallel to EleutherAI’s Pythia, whose absence of character-projection reads as success at being a pure research instrument.
- Reception split (a split in tone, not a documented dispute): a celebratory/playful register (the torrent-heist joke, Together’s “Can you feel the MoE?” copy, general delight in the stunt) coexists with a risk-flagging register (@davidad, 2024-01-23, on open-weight misuse: “could potentially be scaffolded like ChaosGPT by 1 angry teen” — though his “Mixtral-8.6T” figure matches no shipped model, reading as hypothetical extrapolation). Nothing in this corpus shows the two registers arguing directly, so this is noted as a split, not promoted to Contested.
- Longitudinal: GPT-4-is-secretly-MoE rumor → bare magnet link → the format becomes the story → uncensored finetune within a week → “first truly convincing openly licensed implementation” → serving-speed war → agentic substrate → 8x22B repeats the choreography → long tail as an interpretability specimen until the weights quietly stop being hosted.
- tk — whether any janus-sphere character/companion reading of Mixtral itself exists (searched, not found); @Sauers_’s structured base-model comparisons and jd_pressman’s “dunking on 8x22B” / “BigVAE” tweets, known only via excluded retweets — primary URLs not located; the referent of @repligate’s 2024-09-13 “it” (depends on an un-transcribed image)
Records
Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.
Further records
Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.