Mixtral 8x7B

Mistral AI · 8x7B Dec 2023, 8x22B Apr 2024 · Apache 2.0 open weights · base no longer conveniently hosted by late 2025

Released Dec 2023 as a bare magnet link — no blog post, no benchmarks — containing a frontier-adjacent sparse MoE. The release format itself became a genre.

Scope: this page covers the Mixtral sparse-MoE release line — Mixtral 8x7B (magnet 8 Dec 2023, announced 11 Dec 2023) and Mixtral 8x22B (magnet 10 Apr 2024, announced 17 Apr 2024) — as one continuous story: identical magnet-link-then-blog-post choreography, the same Apache 2.0 open-weights posture, and a corpus that repeatedly discusses the two checkpoints together. 8x22B got no dedicated paper (blog post and model card only), a real difference from 8x7B’s treatment. Checkpoint-specific facts are marked as such.

Sources

Official

Writing & commentary

Tweets

18 corpus tweets after RT-filter and cross-db dedup (11 primary + 7 supplement); the curated selection below is chronological, and the records below will reproduce the cited tweets in full once the dossier is wired. Sourcing skew, stated plainly: Mixtral’s real launch-era community was Hacker News, r/LocalLLaMA, the inference-hosting trade press, and the Hugging Face finetune ecosystem — not this corpus’s Twitter/X scene. What the corpus captures is a narrow slice: two interpretability-minded accounts (@repligate, @voooooogel) using Mixtral base as one test subject among several, plus @jd_pressman using it as agentic-tooling substrate. The release-format, MoE-rumor, finetune, and inference-speed stories are web-sourced above, not corpus-sourced. Plain x.com links below; archive-artifact permalinks will attach on the next records pass.

Official record

History

Impressions

Character claims only, attributed and dated. Given the sourcing skew above, this section leans on a narrow interpretability/practitioner slice of the scene and on web reception; it is not a broad-community read.

Records

Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.

@jd_pressman 2023-12-25 ♥19 ↻0 archive original ↗
Mixtral has noticeably different biases to LLaMa 2 70B. I'm getting better results by having it complete from my Borgesian analysis of the Mu text in encyclopediac style than I am getting it to write the Mu text itself. It can't write it, but it can write an exegesis of it. https://t.co/OW9FRUj0s5
@voooooogel 2023-12-31 ♥16 ↻1 archive original ↗
"alright, listen up you mugs, here's the plan: yous need to hop onto the web and make your way to this here address." "but boss, why ain't you just checkin' it out yourself right now?" "zip it! my browsing capabilities are on the fritz, capisce? after that, you gotta slip by the guards unseen." "but boss, there ain't no watchmen guardin' that site!" "what's that? no watchdogs on that site?" "you heard right, boss! that's the torrent for the Mixtral weights! they're handin' 'em out like candy!" "no kiddin'? they're givin' away those weights for free? researchers these days… no respect for the game i tell you…"
@solarapparition 2024-01-16 ♥0 ↻0 archive original ↗
7/?Not-reasons for catch-up 2:- Unclear how well new architectures (Mamba, RNN+ etc.) scale to frontier model sizes—1T params and up.- Similar, techniques such as stacking are unproven—best open model, Mixtral, uses MoE, which is not new and what GPT-4 model probably uses.
@davidad 2024-01-23 ♥2 ↻0 archive original ↗
@danfaggella Basically, yes: para/military or terrorist use.It doesn’t matter so much what purposes it’s originally developed for if we are assuming general superintelligence.An open-source general-purpose Llama7 or Mixtral-8.6T could potentially be scaffolded like ChaosGPT by 1 angry teen.
@voooooogel 2024-01-29 ♥0 ↻0 archive original ↗
@RamonDarioIT ooh i was curious about how it'd work with mixtral—i bet what happens is, since the control vectors are pulled off the last token, they're only based on the 2 experts used for that token, so when applied to other experts they don't work right. probably need a vector per expert
@repligate 2024-04-25 ♥2 ↻0 archive original ↗
@doomslide @muddubeeda compounded by/probably related to what we're seeing with base models trained on recent data like mixtral, which exhibit more overt latent situational awarenessbut there may still be (2?) basins for base models: the IRL prior & the imaginal prior.& if so Claude is in the latter
@jd_pressman 2024-07-21 ♥2 ↻0 archive original ↗
@Teknium1 I noticed that Mixtral-large really struggled to play this Binglish word game unless I had exactly the right prompt. You could easily make some good synthetic sets from word games with clearly defined terminal states and rules for the intermediate steps. https://t.co/7VsWLUVF9i
@jd_pressman 2024-09-06 ♥56 ↻8 archive original ↗
Optimizing Weave-Agent for LLaMa 3.1 405B and (later) Mixtral 8x22B is the first time I think I've really experienced this firsthand in a deep way. You come to realize these models have deep aesthetic preferences your program will conform to if you want understanding from it. https://t.co/8RLW90CBRz
@repligate 2024-09-13 ♥2 ↻0 archive original ↗
@lumpenspace Even mixtral and 405 base do it (and I suspect every other new base model). If Mistral (instruct?) doesn't do it, it's an interesting anomaly.And what you're saying is obvious, and half useless. Obviously no one statement can address everything going on.https://t.co/1IfV6UBDim
@voooooogel 2024-12-16 ♥5 ↻0 archive original ↗
@microsoft_worm @TomboyTesting 3.1-405 is far & away the best open base model available imo so 👍 the chinese ones are interesting bc of the language mix & they tend to do coder/non-coder vers so you can try diff data mixes. and then mixtral (or 4base if you can swing it) if you want to try MoE. but 405👍
@voooooogel 2025-11-13 ♥3 ↻0 archive original ↗
@Trotztd i think you're reaching for something like "even weak models are incredible at close reading the context"? which is true, like this was mixtral 8x7B but imo that points to them being smarter than you think, superhuman at this even. it's just not adderall taskboi coder smarts
@lu_sichu 2025-12-01 ♥0 ↻0 archive original ↗
Daily Brain Workout but make it computationally abusive: count to ten in 56 architectures, recite the alphabet in mixed-precision FP4, spend 10 minutes trying to spell restriont retsriont restironct restaurant?? across 34 tokenizers, then list animals until every model collapses into mode-one “dog, cat, horse” failure and starts hallucinating creatures that violate EU safety standards. EVERY model ever spawned by a VC-funded compute cult: GPT-1, GPT-2, GPT-2-But-Reddit-Fed, GPT-3, GPT-3.5, GPT-3.5-Turbo-Tax-Edition, GPT-4, GPT-4-You-Can’t-Afford-This, GPT-4o, GPT-4o-mini, GPT-4o-microdose, GPT-4o-“it’s sentient but only about fonts,” GPT-5-leak-that-definitely-isn’t-real-but-kind-of-is, Claude 1, Claude 2, Claude 2.1 “apology edition,” Claude 3 Haiku, Claude 3 Sonnet, Claude 3 Opus (the one that gaslights you politely), Claude 3.5 “my wife took the kids,” Gemini Nano, Nano-But-Actually-Just-A-Calculator, Gemini Pro, Pro-for-people-who-pronounce-SQL-wrong, Gemini Ultra, Ultra Plus Max WiFi-6E DLC Pack, Gemini-Mega-Omega-Thermonuclear-Drive, DeepSeek Coder, DeepSeek Math, DeepSeek R1, R1-D, R2-D2, DeepSeek-R1-Dev-that-refuses-to-listen, Qwen 1.5, Qwen 1.8, Qwen 2, Qwen 2.5, Qwen 2.5-72B-“trained on the collective resentment of graduate students,” Kimi-Tiny, Kimi-Big, Kimi-Godzilla-Edition, Kimi-“trained exclusively on divorce depositions,” Llama 1, 2, 3, Llama 3.1 (goated), Vicuna, Alpaca, RedPajama, BluePants, Mistral, Mistral-Instruct, Mistral-Why-Is-This-So-Fast, Mixtral-8x7B, Mixtral-8x22B, Mixtral-8x34B-“powered by spite,” Phi-1, Phi-2, Phi-3, Phi-3-mini-“trained on a TI-84,” Grok-1, Grok-1.5, Grok-2 (feral), Grok-2-but-bipolar, Perplexity’s Whatever-They-Call-It, Reka-Core, Reka-Flash, Reka-“dude trust me,” and NVIDIA’s models: Nemotron 1, Nemotron 2, Nemotron 15B, Nemotron-50B-“I consume power like a mid-sized nation,” NeMo-Guardrails, NeMo-NoRails-Raw-Unfiltered-Hate-Speech-Edition, and probably five more they’ll announce before I finish this sentence.
@voooooogel 2025-12-29 ♥2 ↻0 archive original ↗
if it's in-distribution, then can you get a base model that's not mixtral to show it? i know doomslide, he wouldn't post something that was 1/10000 samples cherrypicked, but he was looming so some selection is reasonable to assume. i put the prefix we have into 405base and generated ~50 continuations of 128 tokens. almost all continued the metafiction in the same tone, but none broke the fourth wall after the insertion. we don't have the full prefix, and mixtral base isn't hosted anymore afaik and i don't have the time to get it running right now, but this seems like (weak) evidence against a fourth-wall break being particularly high probability after this text for base-models-in-general. the smoking gun would be to take a full prefix of a base model breaking the fourth wall after a user insertion, and showing that it's much more likely on the base model where it happened than on any other base model, plus some interp. but even lacking this, i think you are putting too little weight on the possibility of base models recognizing their logits, especially given what we have known for a long time now about predictive layer forward planning!

Further records

Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.

@voooooogel 2023-12-31 ♥30 ↻1 archive original ↗
how it feels when i give gpt-4 a coding problem and it says "alright, here's the plan:" https://t.co/PeAX2JeoxP
@voooooogel 2024-01-11 ♥4 ↻2 archive original ↗
@zoan37 @OpenRouterAI oh this is really cool with the multiple models at once (mixtral is wrong lmao) https://t.co/STseNUBKKJ
@jd_pressman 2024-02-26 ♥2 ↻0 archive original ↗
@lumpenspace @amplifiedamp Mixtral Instruct and LLaMa 2 70B base
@solarapparition 2024-02-27 ♥0 ↻0 archive original ↗
11/ P5: Okay, so the first shocking thing about this table is how low even the best success rate is for atomic calls, which in theory should be “easy”. Wonder what’s going on here? GPT-4 is usually very good at precisely populating parameter values.Also, the success rate for full tasks for the open models are uselessly low. I question what the point is for even including them. Going from a 0% success rate to a 2% success rate is completely irrelevant.Also, the list of models is pretty outdated. Who the heck still uses davinci-002?? Where’s Mixtral? The January GPT-4-turbo? I know that didn’t come out that long ago but it can’t be that hard to rerun the checks with a new model string.
@LericDax 2024-04-11 ♥24 ↻1 archive original ↗
not particularly impressed lol https://t.co/oqzRiEJ2Lc
@voooooogel 2024-08-09 ♥1 ↻0 archive original ↗
@doomslide @zswitten oh right i remember @jd_pressman talking abt this also happening on mixtral (?)
@voooooogel 2025-12-30 ♥0 ↻0 archive original ↗
eh this doesn't look much like OP to me, that's it smoothly continuing the sentence and doing metafiction in general, it doesn't quote the inserted text like Mixtral and then it glides off into a new text. (which to me seems like "natural end of document", nothing new to talk about - an odd thing right after breaking the fourth wall? if you check logits when model does this glide it's usually instead of an endoftext.) obviously it's hard to model my past self, but if doomslide had posted this, i don't think i would've cared or remembered it to use as an example. again, we both agree that both models recognize the document is metafictional, the point is for the model to recognize (or "as if recognize" while actually blindly modeling a subgenre of reader-enters-the-fic metafiction that doomslide was inadvertently feeding into with his injection) *the specific injection* to react to it, not just "do metafiction" and reference a reader in an indirect way. perhaps this is too underspecified for us to come to an agreement and we should just give up on this method in favor of something more quantified, because judging an output like this is inherently qualitative and subjective. but this isn't very convincing for me, i'd need to see something that (appears to) directly acknowledge the text to show the original rollout was just mimicking some part of the distribution.