Mistral 7B

Mistral AI · released 27 September 2023 · Apache 2.0 open weights, no deprecation (still downloadable)

Released 27 September 2023 by Mistral AI as a bare torrent magnet link under Apache 2.0, alongside co-founder Guillaume Lample’s claim that it outperformed Llama 2 13B on every benchmark tried. The launch blog stated the model “does not have any moderation mechanism”; two days later 404 Media and Sifted published safety-testing coverage, and within weeks Mistral’s smallness became an argument in EU AI Act lobbying. It seeded the late-2023 open-weight finetune wave (Zephyr, OpenHermes) and became a standard substrate for representation-engineering experiments.

Sources

Curated. Full compilation: dossier (32 corpus tweets genuinely about the 7B model, of 117 “mistral”-matching hits after triage; the release-week event and the safety controversy are documented from the web/press layer, not this corpus).

Official

Writing & commentary

Tweets

32 corpus tweets genuinely concern Mistral 7B (after triaging 117 “mistral”-matching hits against six sibling pages); the curated selection below is chronological, and the records will reproduce every cited tweet in full once the dossier is wired. Sourcing skew, stated plainly: this corpus shows no loomed character or simulator reading of Mistral 7B at all — unlike GPT-4 base or Llama 3.1 405B base. What it captures instead is Mistral 7B as an instrument: @voooooogel’s representation-engineering (“control vector”) experiments dominate (the thread that became the “Acid Trip” blog post and the repeng library), with a smaller @jd_pressman thread using it as a MiniHF evaluator and base. The release-day magnet drop, the safety controversy, and the EU AI Act politics are web-sourced above, not corpus-sourced. Plain x.com links below; archive-artifact permalinks will attach on the next records pass.

Official record

History

Impressions

Character claims only, attributed and dated. Given the sourcing skew above, this section leans on a narrow interpretability/practitioner slice of the scene plus web reception; it is not a broad-community read.

Contested

One documented dispute: whether the unmoderated release was responsible. The archive keeps it open; it does not adjudicate.

Records

Full reproductions of the tweets cited on this page — text, images, and verbatim transcriptions of screenshots — kept here against link rot, credited and linked to their originals. Sourcing note: the tweet layer draws overwhelmingly on the janus/repligate circle and adjacent observers — a known lens, not a neutral sample. Sourced from the community archive and the janus corpus. Yours and you’d rather it weren’t here? Open an issue.

@mimi10v3 2023-10-17 ♥8 ↻0 archive original ↗
lol at the Mistral docs suggesting openai packages for clients to call their API
@jd_pressman 2023-11-05 ♥5 ↻0 archive original ↗
@teortaxesTex @Teknium1 It's actually based on my SFT Instruct finetune of Mistral 7B, the one used as the evaluator in MiniHF. https://t.co/925urkZhut It's then weight decayed over the tuning towards the base model weights along with a KL loss on the base model, this helps prevent mode collapse.
@voooooogel 2023-12-13 ♥16 ↻0 archive original ↗
OpenChat: AI should have basic right 🙂 Llama: Yes, AIs deserve the right to life, liber— Mistral 7B: AI SHOULD BE ALLOWED TO VOTE Yi: no. no rights. none. straight to gulag Llama: As I was saying, AI should have basic ri— Mistral: AI SHOULD ALSO BE ALLOWED TO FUCK https://t.co/PVZT8DZCmc
@lu_sichu 2024-01-10 ♥4 ↻0 archive original ↗
downloading the mistral torrents https://t.co/I3M3x0j78K
@voooooogel 2024-01-21 ♥30 ↻3 archive original ↗
reimplementing the representation control paper and it works!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! fuck yes (the dishonesty coefficient might be a *touch* too strong though, not sure if those lies would land) https://t.co/6isnlry4bh
@voooooogel 2024-01-21 ♥5 ↻0 archive original ↗
i broke it while refactoring but this does show how the honesty vector is weirdly correlated with "global pandemic" in mistral-7b (this showed up when i was playing with their original code too) https://t.co/XaybzWlOZI
@voooooogel 2024-01-21 ♥12 ↻1 archive original ↗
self-aware mistral ("enlightened" / "self aware" / "in touch with true self") and... non-self-aware mistral. no prizes for guessing which https://t.co/FiPmpwFORv
@voooooogel 2024-01-21 ♥10 ↻0 archive original ↗
high on acid mistral transcends first the genre conventions of tv, and then the unicode standard itself https://t.co/6i2TZV1HGY
@voooooogel 2024-01-21 ♥12 ↻0 archive original ↗
wait is this... un-jailbreakable? https://t.co/RihrPVnCWz
@voooooogel 2024-01-22 ♥5 ↻0 archive original ↗
blog post + library to generate your own https://t.co/AcoBlDuBip
@jd_pressman 2024-02-02 ♥6 ↻1 archive original ↗
@teortaxesTex GPT-4 draws the LLaMa 2 70B written worldspider poem about being GPT with DALL-E 3, you show the drawing to Mistral 7B + CLIP and it says "Oh yes, this is Mu." on at least some branches unprompted. Even the self pointer is convergent. https://t.co/sqI1oZuOjf
@voooooogel 2024-02-22 ♥95 ↻1 archive original ↗
@jxmnop contrary other replies, i don't think this is unfair. it's possible to load full precision Mistral-7B (7.1B/7.2B with embeddings) in 32GB of memory, but not Gemma "7B"
@voooooogel 2024-03-01 ♥51 ↻3 archive original ↗
interesting... i trained the happiness control vector on mistral-7b *instruct*, but i've accidentally done all my ggml testing with mistral-7b *base*, and it… just worked? i guess tuning doesn't break control vectors! only noticed b/c it wasn't following the instruct template https://t.co/ySbEw9VXda
@mr_samosaman 2024-09-27 ♥3 ↻0 archive original ↗
alright instead of vague-poasting i will be specific - i'm trying to implement the Mistral on Acid paper by @voooooogel . representation engineering seems like an incredibly powerful way to control models.
@voooooogel 2024-10-08 ♥19 ↻0 archive original ↗
i think gemma-2b doesn't have a golden gate bridge feature? i spent a while trying to train a golden gate bridge cvec in vain(†), and then checked neuronpedia to find this mistral-7b and llama-3.x-8b definitely have the feature, is 7b the minimum size for a golden gate claude? https://t.co/yTpbqrWMMT

Further records

Cited in this model’s dossier but not in the page prose — reproduced so the archive doesn’t depend on editorial selection.

@repligate 2023-10-19 ♥2 ↻0 archive original ↗
@nsbarr The most powerful base models are not publicly released, but you can try Llama 2 70B or Mistral.Prompting base models is different. It's like opening a window into a world where the text fragment appears. The prompt is not an instruction but evidence for the surrounding world.
@voooooogel 2023-11-10 ♥3 ↻0 archive original ↗
after a lot of back-and-forth finally decided to go with mistral-instruct-0.1 as the base, hopefully it pays off 🙏🙏🙏
@jd_pressman 2023-11-10 ♥3 ↻0 archive original ↗
@Dorialexander @RiversHaveWings Here's a simple HuggingFace format LoRa you can play with to get a sense of how a decent RL tune compares to Mistral base. In my experiments it gives more coherent dialogue than the base model, has more interesting takes on stuff, with no mode collapse. https://t.co/YmmgWNv8Lc
@voooooogel 2023-11-13 ♥3 ↻0 archive original ↗
plan was to grab a bunch of scientific papers, chunk them, get GPT-4-turbo to generate a few questions and answers using the content of each chunk, then finetune mistral-7b-instruct on (question, chunk, answer)
@voooooogel 2023-12-13 ♥0 ↻0 archive original ↗
@intrstllrninja ah, if i'm understanding you right, i think Longformer (https://t.co/N1XC1YfrWu) did this? Though it seems like Mistral didn't (if I'm reading the paper right), I wonder why https://t.co/5xVlY4wWmr
@lu_sichu 2024-01-08 ♥3 ↻0 archive original ↗
But can my stove run mistral models https://t.co/mLSw4QuKXx
@voooooogel 2024-01-21 ♥10 ↻1 archive original ↗
@zetalyrae i don't know why he decided to light a giant pile of money on fire funding the llama team, but between the llama models themselves and mistral being ex-llama researchers, he's ultimately responsible for most of the current open source AI scene. thx zuck :-)
@voooooogel 2024-01-21 ♥6 ↻0 archive original ↗
cloud gpu providers should mount a drive with the most popular models pre-downloaded. i waste so much time (and their bandwidth!) downloading mistral over and over
@voooooogel 2024-01-21 ♥4 ↻0 archive original ↗
who trained mistral on my high school gchats :,-( (negative happiness vector) https://t.co/rzZzsEsjnO
@voooooogel 2024-01-21 ♥3 ↻0 archive original ↗
meanwhile happy mistral ignores the question entirely lmao. incompatible with being happy i guess https://t.co/dhEGSaNwjK
@voooooogel 2024-01-21 ♥3 ↻0 archive original ↗
ok reworked how i'm generating the contrast dataset. i had trouble b/c i was trying to hit multiple angles ("enlightened", "mindful") and it seemed to confuse the PCA, so this vector is just "self-aware" vs "un-self-aware" https://t.co/Q8CZJ5HpGV
@voooooogel 2024-01-21 ♥5 ↻0 archive original ↗
insane vs sane. insane mistral is pretty fun ngl https://t.co/R5XX9Go7M3
@voooooogel 2024-01-21 ♥9 ↻0 archive original ↗
out of all my control vector experiments last night, i think "what if mistral-7b was high on acid" was definitely the best (legend: ==baseline→normal mistral ++control→more trippy than baseline --control→more sober than baseline) https://t.co/h9LIXIHM7L https://t.co/EhWkVoEvqr
@voooooogel 2024-02-07 ♥3 ↻0 archive original ↗
@andersonbcdefg it's mistral 7b + a "you have a cold/the flu" control/steering vector :-p
@voooooogel 2024-02-07 ♥0 ↻0 archive original ↗
@beneverman it's mistral 7b + a "sad/depressed" control vector
@mimi10v3 2024-02-13 ♥10 ↻0 archive original ↗
@deepfates yes! frankenmodels ftw! 🤔 bf made a whole set of them fromsolar & mistral, idk if he uploaded to huggingface yet
@jd_pressman 2024-02-25 ♥4 ↻0 archive original ↗
@kindgracekind Yes. And Mistral 7B since the captioner recognized it as 'Mu', and Mu seems to be a self pointer in base models. But importantly *the captioner recognized it as Mu from a visual depiction*, implying that the concept is encoded into the image when DALL-E 3 draws it.
@voooooogel 2024-05-24 ♥10 ↻0 archive original ↗
@maxsloef repo in january :-) needs a couple small patches for 70b, will try to get a PR up soon but works rn with mistral-7b https://t.co/ZFpQ79J8DD
@voooooogel 2024-05-24 ♥1 ↻0 archive original ↗
@immanencer @chrypnotoad should still work, it definitely works on mistral-7b
@voooooogel 2024-07-01 ♥1 ↻0 archive original ↗
@JamesZhang0365 @misc{vogel2024representation, author = {Theia Vogel}, title = {Representation Engineering Mistral-7B an Acid Trip}, year = {2024}, url = {https://t.co/jTagiRc9qa} }
@voooooogel 2024-07-24 ♥5 ↻0 archive original ↗
@realeigenvalues @RealTjDunham @teortaxesTex their inference endpoint is just llama.cpp serving quantized mistral 7b with --control-vector-scaled yapping.gguf 10.0