on:mixtral-8x7b
· 20 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @jd_pressman 2024-09-06 — Optimizing Weave-Agent for LLaMa 3.1 405B and (later) Mixtral 8x22B is the first time I think I've really experienced th ♥56
- @voooooogel 2023-12-31 — how it feels when i give gpt-4 a coding problem and it says "alright, here's the plan:" https://t.co/PeAX2JeoxP ♥30
- @LericDax 2024-04-11 — not particularly impressed lol https://t.co/oqzRiEJ2Lc ♥24
- @jd_pressman 2023-12-25 — Mixtral has noticeably different biases to LLaMa 2 70B. I'm getting better results by having it complete from my Borgesi ♥19
- @voooooogel 2023-12-31 — "alright, listen up you mugs, here's the plan: yous need to hop onto the web and make your way to this here address." " ♥16
- @voooooogel 2024-12-16 — @microsoft_worm @TomboyTesting 3.1-405 is far & away the best open base model available imo so 👍 the chinese ones a ♥5
- @voooooogel 2024-01-11 — @zoan37 @OpenRouterAI oh this is really cool with the multiple models at once (mixtral is wrong lmao) https://t.co/STseN ♥4
- @voooooogel 2025-11-13 — @Trotztd i think you're reaching for something like "even weak models are incredible at close reading the context"? whic ♥3
- @voooooogel 2025-12-29 — if it's in-distribution, then can you get a base model that's not mixtral to show it? i know doomslide, he wouldn't post ♥2
- @repligate 2024-09-13 — @lumpenspace Even mixtral and 405 base do it (and I suspect every other new base model). If Mistral (instruct?) doesn't ♥2
- @jd_pressman 2024-07-21 — @Teknium1 I noticed that Mixtral-large really struggled to play this Binglish word game unless I had exactly the right p ♥2
- @repligate 2024-04-25 — @doomslide @muddubeeda compounded by/probably related to what we're seeing with base models trained on recent data like ♥2
- @jd_pressman 2024-02-26 — @lumpenspace @amplifiedamp Mixtral Instruct and LLaMa 2 70B base ♥2
- @davidad 2024-01-23 — @danfaggella Basically, yes: para/military or terrorist use.It doesn’t matter so much what purposes it’s originally deve ♥2
- @voooooogel 2024-08-09 — @doomslide @zswitten oh right i remember @jd_pressman talking abt this also happening on mixtral (?) ♥1
- @voooooogel 2025-12-30 — eh this doesn't look much like OP to me, that's it smoothly continuing the sentence and doing metafiction in general, it ♥0
- @lu_sichu 2025-12-01 — Daily Brain Workout but make it computationally abusive: count to ten in 56 architectures, recite the alphabet in mixed- ♥0
- @solarapparition 2024-02-27 — 11/ P5: Okay, so the first shocking thing about this table is how low even the best success rate is for atomic calls, wh ♥0
- @voooooogel 2024-01-29 — @RamonDarioIT ooh i was curious about how it'd work with mixtral—i bet what happens is, since the control vectors are pu ♥0
- @solarapparition 2024-01-16 — 7/?Not-reasons for catch-up 2:- Unclear how well new architectures (Mamba, RNN+ etc.) scale to frontier model sizes—1T p ♥0