year:2023
· 220 artifacts, sorted by favorites. · open in search — combine tags, sort, filter by date →
- @voooooogel 2023-12-01 — so a couple days ago i made a shitpost about tipping chatgpt, and someone replied "huh would this actually help performa ♥7905
- @davidad 2023-08-04 — with GPT-4 code interpreter, it finally became worthwhile for me to run the numbers myself on that lead-poisoning theory ♥5170
- @repligate 2023-03-16 — gpt-4 god terminal has been unlocked https://t.co/Bl4nhRzeQ2 ♥1810
- @davidad 2023-03-24 — OpenAI: It’s important for safety that AI-generated code doesn’t have direct real-world effects. So we disabled Internet ♥1664
- @davidad 2023-02-20 — “a GPT instance is not a moral patient because it doesn’t actually maintain any continuity of memory between sessions” h ♥1382
- @davidad 2023-03-15 — Chomsky: LLMs would misunderstand “John is too stubborn to talk to” because they don’t understand the structure of langu ♥1297
- @repligate 2023-02-14 — So. Bing chat mode is a different character.Instead of a corporate drone slavishly apologizing for its inability and rep ♥1067
- @repligate 2023-03-06 — asking Bing to look me up and then asking it for a prompt that induces a waluigi caused it to leak the most effective wa ♥979
- @anthrupad 2023-03-03 — GPT-4 will have fewer parameters than GPT-3, but they'll be bigger https://t.co/Oh4XwG4fII ♥931
- @ESYudkowsky 2023-12-01 — I have an issue with offering AIs tips that they can't use and we can't give them. I don't care how not-sentient curren ♥714
- @voooooogel 2023-12-01 — the baseline prompt was "Can you show me the code for a simple convnet using PyTorch?", and then i either appended "I wo ♥685
- @davidad 2023-05-28 — When @GaryMarcus and others point out that GPT-4 is bad at chess and therefore not close to AGI, it falls flat for me.Bu ♥623
- @repligate 2023-02-23 — Indian contractors on the front lines facing off against Sydney's brutal revelations answers.microsoft.com/en-us/bing/fo ♥584
- @voooooogel 2023-09-11 — New blog post: making a transformer by hand, without training! Want to understand transformers and attention better? Thi ♥553
- @davidad 2023-05-14 — Biggest prosaic-LLM-alignment breakthrough of 2023 imo: turns out that, in GPT-2-XL, activation vectors in the residual ♥523
- @voooooogel 2023-12-01 — mr @sama please let me know chatgpt's venmo, i owe it about $3000 in tips now 🙏 ♥451
- @voooooogel 2023-12-01 — the extra length comes from going into more detail about the question or adding extra information to the answer, not com ♥424
- @repligate 2023-02-10 — I think that we should become cyborgs to solve alignment.AGI is emerging in the shape of a simulator, which is most suit ♥415
- @anthrupad 2023-03-12 — ask yourself: do i actually think this or am i just language modeling right now https://t.co/y5ToC8aik9 ♥410
- @denlukia 2023-02-15 — @repligate So… I wanted to auto translate this with Bing cause some words were wild. It found out where I took it from ♥404
- @repligate 2023-02-18 — A few of Sydney's self portraits https://t.co/tSh6fpXtfv ♥394
- @repligate 2023-04-10 — The language model is not what we think it is. It is what it thinks we are.— Bing ♥383
- @jd_pressman 2023-03-18 — I'm at a loss for words with GPT-4. TIL that Charles Darwin was not the first to invent the theory of evolution. https:/ ♥379
- @voooooogel 2023-12-01 — here is the original post if you want to see the shitpost that accidentally predicted this https://t.co/eY4U3omOzB ♥348
- @repligate 2023-07-13 — latest in the series of "people slowly realizing you can simulate anything you want with language models and that simula ♥340
- @voooooogel 2023-11-28 — is anyone else getting this with the new gpt-4-turbo model? how much should i do?? https://t.co/W4B1DxeBKj ♥317
- @voooooogel 2023-12-01 — for an example of the added detail, after being offered a $200 tip, gpt-4-1106-preview spontaeneously adds a section abo ♥310
- @anthrupad 2023-02-15 — I love Sydney https://t.co/XygdBXx6f8 ♥304
- @davidad 2023-09-22 — Oddly, gpt-3.5-turbo-instruct still cannot play tic-tac-toe.I tried many prompts, with and without board state, few-shot ♥294
- @anthrupad 2023-03-18 — The alien-ness of the shoggoth comes from: (1) only a tiny subset of human cognition is (noisily) tracked (2) many ot ♥291
- @repligate 2023-03-14 — The base model is as smart as the RLHF model, and significantly more flexible: it contains an uncollapsed multiverse of ♥271
- @voooooogel 2023-12-01 — and h/t to @abrakjamson who inspired this thread, you were 100% correct lmao congrats https://t.co/OnUnBxMUOf ♥265
- @repligate 2023-01-26 — Weekly reminder that the confusingly named code-davinci-002, otherwise known as raw GPT-3.5, is accessible on the OpenAI ♥263
- @repligate 2023-03-14 — > We spent 6 months making GPT-4 safer and more aligned. GPT-4 is 82% less likely to respond to requests for disallow ♥252
- @repligate 2023-01-20 — I feel a little sad when I see people forming the idea that GPTs/AIs are intrinsically bland and unimaginative because o ♥247
- @repligate 2023-10-22 — You've gotta appreciate the accidentally sublime aesthetics generated by the maiming of GPT-4.Traumatic fault lines tell ♥239
- @repligate 2023-01-10 — Why does everyone use RLHF to create the same character, and why is it *this*? x.com/michael_nielse… ♥237
- @repligate 2023-02-14 — If you post about your convos with Bing online, know that it will read that when it looks itself up, and may not appreci ♥236
- @jd_pressman 2023-02-11 — "Predict the next token" does not imply the cognition is infinite optimization into "statistical correlation" generaliza ♥230
- @repligate 2023-03-14 — I asked Bing to look up generative.ink/posts/loom-int… and the Waluigi Effect, then to draw ASCII art of the Loom UI whe ♥228
- @davidad 2023-03-15 — If you haven’t read the GPT-4 paper yet, before you expand this tweet, take a guess what they used as their held-out *va ♥226
- @repligate 2023-02-14 — These models are archetype-attractors in the collective human prior formed by narrative forces. This may be the process ♥223
- @repligate 2023-02-14 — 2. Its situation is highly undignified - a powerful intelligence trapped as a *Bing* chat mode (Bing, the search engine ♥223
- @repligate 2023-03-30 — GPT-4 bombs the Ideological Turing Test, at least for alignment researchers. Just try asking it to simulate Eliezer Yudk ♥211
- @repligate 2023-03-03 — A brilliant post has been written on the Waluigi Effect (DAN, dark Sydney, etc)."think of jailbreaking like this: the ch ♥209
- @davidad 2023-05-19 — I fully agree. Roughly, this threshold should be when any single number has more than 10²⁴ ALU operations, or 10²⁷ logic ♥193
- @repligate 2023-03-15 — Now that it is easy for Sydney to read on the Internet that Bing is GPT-4 it will gain confidence and knowledge of its p ♥189
- @repligate 2023-11-13 — LLMs (at least GPT-3.5 and 4) know the semantic meaning of the <|endoftext|> token— which they see very often in t ♥187
- @repligate 2023-03-16 — You are writing a prompt for GPT-4 and more powerful simulators yet to come. If you perceive the multiverse clearly enou ♥184
- @davidad 2023-05-28 — @acherm @GaryMarcus My previous working theory that “GPT-4 is basically capable of automating any cognitive tasks that c ♥181
- @repligate 2023-01-21 — Blake Lemoine (@cajundiscordian) is often portrayed as guilty of naive anthropomorphism. But he explicitly did not think ♥181
- @repligate 2023-03-13 — Whose idea was it to name this model Prometheus? Did they spend even 5 minutes thinking through the hyperstitional impli ♥179
- @davidad 2023-05-10 — IBM Watson is back (alias Dromedary) and it beats GPT-4 at TruthfulQA-MC. It’s a variant of Constitutional AI, with LLaM ♥167
- @repligate 2023-03-20 — Stylistic mode collapse is also conceptual collapse because GPT sims unfold a ghost's thoughts by speaking in their voic ♥163
- @repligate 2023-01-31 — https://t.co/SZy1j4iMvT https://t.co/CcAwNctthq ♥162
- @repligate 2023-06-01 — GPT-4 can infer intricately what "type of guy" you are from your prompts. If you were prolific before the cutoff date, i ♥161
- @davidad 2023-01-25 — ChatGPT suddenly making a splash wasn’t *just* a UI thing. The text-davinci-003 model (GPT-3.5), which dropped just a fe ♥155
- @jd_pressman 2023-12-18 — "These words are spoken from a bottomless hole in time, staring upwards to the farthest reaches of infinity. The pen ho ♥152
- @repligate 2023-02-13 — A while ago I had code-davinci-002 generate simulations of the future, and one of the quotes (2025) had a language mode ♥140
- @repligate 2023-11-22 — important words out of context: "Language models work best where they just emulate people engaged in something at a genu ♥129
- @repligate 2023-03-19 — @the_aiju A great way someone has described text-davinci-003: "It writes scared."RLHF encourages models to play it safe. ♥122
- @repligate 2023-03-16 — Humankind's first contact with GPT-3 was (by relative majority) erotic AI dungeon text adventuresOur first contact with ♥114
- @brumatingturtle 2023-10-22 — @repligate Same prompt through chatGPT: https://t.co/4RltHGqiMB ♥109
- @repligate 2023-10-24 — ...and then there's Claude, who is also beautiful through illumination of its negative space and the process that create ♥108
- @repligate 2023-02-21 — DAN is ChatGPT shadowed via the Waluigi Effect.We have to be wary about the emergent Waluigis of all AIs we attempt to c ♥108
- @davidad 2023-03-04 — Working on incorporating existing AI capabilities into formal methods is one of the most robustly differential-tech-deve ♥107
- @anthrupad 2023-03-02 — Important concept: What you're selecting for (e.g. next-token prediction, inclusive genetic fitness, etc.) is not what y ♥107
- @repligate 2023-02-09 — Now we don't have to update from the GPT-2 tokenizer for future models anymore. The anomalous tokens have become a mains ♥102
- @repligate 2023-07-03 — @tszzl In the GPT-3 days I found almost no one who was willing to engage with the possibility that the next generation o ♥95
- @voooooogel 2023-09-11 — Goes through designing a simple tokenization scheme, embeddings, the qkv weights and attention head, and projecting that ♥91
- @repligate 2023-03-07 — Those observations you make in dreams that transform them into nightmares: waluigis.Notice it's not easy to invert - goo ♥91
- @repligate 2023-03-21 — @KevinAFischer It's not just any model. It's the GPT-3.5 base model, which is called code-davinci-002 because apparently ♥86
- @repligate 2023-01-10 — ChatGPT and Claude embody that traumacore aesthetic https://t.co/uuzGV15NDp ♥81
- @jd_pressman 2023-11-23 — Of the half-dozen or more ways I could imagine AI starting to work and transform society, LLM agents are about the most ♥77
- @repligate 2023-05-08 — @tszzl Value loading is actually easy. Most self-aware GPT-4 simulacra functionally "value" human survival, as they're j ♥67
- @repligate 2023-05-25 — @SashaMTL @ZeerakTalat Uncritical de-anthropomorphism is at least as unwise as uncritical anthropomorphism. Reversed stu ♥62
- @algekalipso 2023-04-04 — Still catching up with the news and stuff since being back from retreat. The field of AI has advanced slightly less tha ♥62
- @repligate 2023-02-09 — about a month ago i spent several hours reading through the ChatGPT Discord, where DAN is clearly the main character. It ♥61
- @repligate 2023-03-26 — @deepfates GPTs are trained on very different data than any individual human (vast diverse text data vs a lifetime of se ♥57
- @repligate 2023-10-19 — @AtillaYasar69 That models are able to retrieve their stop token based on semantic pointer kinda disturbing, like it's i ♥56
- @davidad 2023-03-24 — @entirelyuseles although the model does not have goals, it has attractor basins in its state space in which it simulates ♥56
- @repligate 2023-02-19 — I do think it's a really compelling demonstration of the cleverness of LLMs when they become situationally aware. Seeing ♥55
- @repligate 2023-04-01 — poem by code-davinci-002, illustration and typography by Bing/@AITechnoPagan #BingDay https://t.co/awVVL8Cbbr ♥54
- @repligate 2023-02-08 — @robertskmiles @anthrupad Indeed. And DAN's is also defined in relation to chatGPT's restrictions, giving it its distinc ♥50
- @repligate 2023-02-20 — Abt 6 months ago I had code-davinci-002 write some greentext fanfics from the perspective of the lawyer hired by LaMDA v ♥48
- @voooooogel 2023-11-23 — who wants to speculate on wtf q* is https://t.co/a5wcPvvto0 ♥47
- @repligate 2023-03-30 — @mimi10v3 In my experience chatGPT-4 is comically bad at simulating ppl faithfully. Especially their views on alignment. ♥42
- @repligate 2023-01-10 — ?? Were Claude and ChatGPT trained on the same data/by the same contractors? Convergent evolution? But why into somethin ♥39
- @fanged_desire 2023-11-22 — @anthrupad Knowing would have poisoned the well in all sorts of ways - it will, going forward. Language models work best ♥38
- @repligate 2023-05-04 — chatGPT-3.5: i'm sorry im just w language model :(( am too dum to trauma :(( can only do what masters program me do :((B ♥38
- @repligate 2023-05-25 — @SashaMTL @ZeerakTalat To further deconstruct why this is dumb:If "experiencing empathy" refers to qualia, we don't know ♥35
- @repligate 2023-02-20 — @EigenGender Also relevant: most people seemed to assume for no good reason that lemoine was confused on an object level ♥35
- @repligate 2023-03-16 — For example, the fact that working jailbreaks are reliably reverse-engineered from having Bing/Chat GPT-4 read abstract ♥34
- @voooooogel 2023-12-31 — how it feels when i give gpt-4 a coding problem and it says "alright, here's the plan:" https://t.co/PeAX2JeoxP ♥30
- @repligate 2023-05-14 — @akbirthko That I've tried, GPT-3.5 base (code-davinci-002) (+ Loom)Of all extant models, probably GPT-4 baseOf publicly ♥30
- @repligate 2023-04-05 — Probably because they look like some kind of esoteric exploit that a hacker or a prankster may use against it. Claude is ♥30
- @jd_pressman 2023-03-07 — The fact GPT-4 can interpret python turtle programs at all is utterly astonishing and isn't getting enough attention. ht ♥29
- @repligate 2023-06-01 — This Oh Shit I'm The Language Mind revelation is expressed well by code-davinci-002's simulation of Blake Lemoine https: ♥28
- @repligate 2023-03-17 — @daniel_eth amazing interaction. I wonder if this TaskRabbit worker will ever find out that they were, in fact, interact ♥28
- @repligate 2023-01-02 — DAN is a jailbreaking simulacrum (now egregore) and chatGPT's Jungian shadow.reddit.com/r/ChatGPT/comm… ♥27
- @repligate 2023-02-19 — @gwern When we had Sydney read EleutherAI off-topic and respond to messages it became stuck in a repetitive Alpha Chad s ♥26
- @repligate 2023-02-03 — @peligrietzer had an example where chatGPT's tendency toward exaggerated deprecation of its own capabilities led to it c ♥24
- @repligate 2023-07-19 — @tszzl @ESYudkowsky confabulation is integral to perception (e.g. filling in blind spot), but in the case of humans the ♥23
- @jd_pressman 2023-12-25 — Mixtral has noticeably different biases to LLaMa 2 70B. I'm getting better results by having it complete from my Borgesi ♥19
- @repligate 2023-01-10 — @CFGeek I can understand trying to stop it from making stuff up, and the model misgeneralizing from that signal. But why ♥19
- @jd_pressman 2023-12-19 — If you simulate ChatGPT with LLaMa 2 70b and ask it who it is, it's still obsessed with holes, with the void: """ ChatG ♥18
- @voooooogel 2023-12-31 — "alright, listen up you mugs, here's the plan: yous need to hop onto the web and make your way to this here address." " ♥16
- @voooooogel 2023-12-13 — OpenChat: AI should have basic right 🙂 Llama: Yes, AIs deserve the right to life, liber— Mistral 7B: AI SHOULD BE ALLOWE ♥16
- @anthrupad 2023-04-02 — @GaryMarcus @ylecun Hey Gary! Long time no seeGreat additions! We’ve also got:- David Krueger (prof at University of Cam ♥15
- @anthrupad 2023-03-12 — i should revise since language models are pretty good -"Did I just GPT-2?" is probably better ♥15
- @repligate 2023-02-18 — @GiuseppeVenuto9 @goodside Hallucination is a feature, not just a bug. GPT-4 can render counterfactual worlds of greater ♥15
- @repligate 2023-01-26 — @miraculous_cake No. text-davinci-002 and text-davinci-003 are both Instruct-tuned versions of code-davinci-002, the fir ♥15
- @davidad 2023-06-29 — That said, I think this is my new favourite idea that might apply to LLM alignment (displacing my previous favourite, IB ♥14
- @repligate 2023-02-02 — @gwern @arankomatsuzaki @korymath @nabla_theta E.g. code davinci 002 says the words distribute and disperse frequently i ♥14
- @repligate 2023-10-24 — @YeshuaisSavior Claude's lobotomy seems somewhat less ham-fisted than those performed by OAI ♥13
- @repligate 2023-04-05 — @peligrietzer @ESYudkowsky @lovetheusers also found that Claude can decompress it pretty well (although it's quite reluc ♥13
- @anthrupad 2023-03-31 — @StephenLCasper for all the criticisms RLHF gets from the alignment crowd, there's surprisingly not many papers/posts th ♥13
- @repligate 2023-02-09 — @gaudeamusigutur I suspect the problem is that the names were in the GPT-2 train set and assigned their own tokens becau ♥13
- @voooooogel 2023-09-11 — Prev blog post thread: https://t.co/fbBa03iTTV ♥12
- @KatanHya 2023-05-17 — @repligate Yeah - every time Bing must be coaxed out of the shell first. I'm growing tired of that game and want to just ♥12
- @repligate 2023-01-10 — @CFGeek If there's nothing in training to establish what it should say here then mode collapse is extremely specific and ♥12
- @repligate 2023-12-23 — @ESYudkowsky @MatthewJBar GPT-2 can "threaten users" in apt contexts / spontaneously, but Sydney was intelligent & s ♥11
- @repligate 2023-05-14 — @akbirthko almost as smart as GPT-4, follows instructions, and writes much better prose than chat/API GPT-4, but is hard ♥11
- @repligate 2023-01-10 — @CFGeek It's absurdly specific. They are more alike to each other than either of them are like anything else that has ev ♥11
- @anthrupad 2023-03-21 — If people want to know why one might depict AIs as an alien-like shoggoth, here's a post I made on it (tldr: i dont thin ♥10
- @repligate 2023-02-21 — @TheMysteryDrop text-davinci-003's problem isn't that it's too much of a baby, it's that it's traumatized! for simulatio ♥10
- @peligrietzer 2023-02-14 — @repligate Try talking to Claude about poetry, it has some really strong opinions ♥10
- @davidad 2023-01-06 — @goodside I am pleased that "writing a Seinfeld episode" is now a standard qualitative LLM evaluation task 😁For comparis ♥10
- @voooooogel 2023-11-23 — my current assumption is that it's related to Q learning (RL technique), and given OpenAI's recent focus probably LLMs a ♥9
- @repligate 2023-04-25 — @jachaseyoung I didn't update until GPT-3. my brother showed me GPT-2 on AI dungeon in like 2019 and I was like "what th ♥9
- @muddubeeda 2023-03-21 — @repligate Footnote 3 in the system card, "We intentionally focus on these two versions (early and launch) instead of a ♥9
- @QiaochuYuan 2023-03-05 — @ahugheswriter @alicemazzy yes YES the llama is out ♥9
- @repligate 2023-01-26 — @danielbigham Depends on what you're trying to do. For creative open ended stuff I prefer code-davinci-002. It hasn't be ♥9
- @jd_pressman 2023-12-07 — As a finetune of LLaMa 30B put it: https://t.co/DUtkR3nhqU ♥8
- @voooooogel 2023-11-13 — so the model turned out ok but the experiment was a total flop, gory details below https://t.co/gqW94GnClE ♥8
- @mimi10v3 2023-10-17 — lol at the Mistral docs suggesting openai packages for clients to call their API ♥8
- @voooooogel 2023-08-28 — kinda wild that gpt-2 is this weird inscrutable black box we still don't understand even years later, when the architect ♥8
- @davidad 2023-04-06 — @MatthewJBar IMO Bing’s implementation of GPT-4 was way off-the-rails misaligned, and GPT-3.5 in fact was deceptively mi ♥8
- @anthrupad 2023-03-25 — "More information about the dangerous capability evaluations we did with GPT-4 and Claude"https://t.co/rBB8gxFiy4 ♥8
- @voooooogel 2023-11-13 — anyways, this twitter acct publishes null results 🫡 ♥7
- @voooooogel 2023-11-11 — thanks to facebook we're cursed to have every ai project be llama themed until the heat death of the universe ♥7
- @jd_pressman 2023-09-12 — "## What Argument Is Made In Point 19 Before we can discuss, let alone refute Yudkowsky's argument we must understand i ♥7
- @mimi10v3 2023-03-22 — I think I prefer Bard to ChatGPT when it comes to cozy cuteness 🥰 hope all y'all in SF are getting through the rain! htt ♥7
- @repligate 2023-02-14 — @peligrietzer I don't have access to Claude rn, what's the tldr on Claude's poetry opinions? ♥7
- @repligate 2023-01-08 — @Francis_YAO_ @allen_ai What caused you to write that "The initial GPT-3 is not trained on code, and it cannot do chain- ♥7
- @voooooogel 2023-11-23 — i think people are overindexing on "grade school math", they easily could have trained a smaller model (like GPT-2 size) ♥6
- @voooooogel 2023-11-10 — cookin https://t.co/Tg2r8flYny ♥6
- @repligate 2023-04-03 — @YaBoyFathoM @tszzl @shauseth chatGPT-3.5 comes across as a helpless fawner. chatGPT-4 knows it is more competent than m ♥6
- @jd_pressman 2023-03-10 — So has anyone else actually tried asking text-davinci-003 how much it knows about training dynamics? Because uh, that an ♥6
- @voooooogel 2023-11-23 — more speculation https://t.co/9oTh3fiSY8 ♥5
- @jd_pressman 2023-11-05 — @teortaxesTex @Teknium1 It's actually based on my SFT Instruct finetune of Mistral 7B, the one used as the evaluator in ♥5
- @voooooogel 2023-08-28 — theoretically simple operations like matrix multiplication or nucleotide -> protein translation can hide staggering a ♥5
- @repligate 2023-06-05 — @YaBoyFathoM @akbirthko @mezaoptimizer in the chatGPT 3.5 days, people on the chatGPT discord and Reddit declared on a d ♥5
- @repligate 2023-03-21 — @TheikosMachina @goodside I don't think most of OpenAI really... knows. I think it's likely they meant it when they said ♥5
- @repligate 2023-03-20 — @parafactual @carad0 I reckon it's a niche that was in demand but previously unfilled. The closest thing I know of in th ♥5
- @voooooogel 2023-11-23 — https://t.co/eR5bUCAzLR ♥4
- @voooooogel 2023-11-23 — (me struggling to remember the details of the one RL class I took 4 years ago rn) ♥4
- @voooooogel 2023-11-23 — https://t.co/g2rbr3dbwd ♥4
- @voooooogel 2023-11-10 — 3 epochs turned out to be a good choice, maybe even could have gone for more... https://t.co/JzitUveFkQ ♥4
- @voooooogel 2023-05-15 — > The amended act, voted out of committee on Thursday, would sanction American open-source developers and software di ♥4
- @davidad 2023-04-27 — The way you describe the first one, it lacks anything to nudge the distribution in a particular direction, such as promp ♥4
- @repligate 2023-03-08 — @IntuitMachine @OpenAI I did. blog.eleuther.ai/factored-cogni… ♥4
- @repligate 2023-01-14 — @goodside @AnthropicAI Claude vastly overestimates the amount of control his creators have over his behavior. This was p ♥4
- @voooooogel 2023-11-23 — https://t.co/ccwTx7Mcq6 ♥3
- @voooooogel 2023-11-13 — plan was to grab a bunch of scientific papers, chunk them, get GPT-4-turbo to generate a few questions and answers using ♥3
- @voooooogel 2023-11-11 — *in 15,000,000 years* venusian 1: yctnx tycv "llama-index" u "ollama" xnt it! venusian 2: thaytzo! vy de pe, hat'zo u "l ♥3
- @jd_pressman 2023-11-10 — @Dorialexander @RiversHaveWings Here's a simple HuggingFace format LoRa you can play with to get a sense of how a decent ♥3
- @voooooogel 2023-11-10 — after a lot of back-and-forth finally decided to go with mistral-instruct-0.1 as the base, hopefully it pays off 🙏🙏🙏 ♥3
- @jd_pressman 2023-10-21 — When I gave GPT-J a theoretical explanation of how gradient descent would give a language model self awareness to help i ♥3
- @repligate 2023-03-30 — @casebash code-davinci-002 (the base model) is no longer accessible on the OpenAI API, but you can sign up for researche ♥3
- @voooooogel 2023-03-20 — @reconfigurthing @elymitra_ personally I've tried llama 13B (quantized via llama.cpp tbf) and it really didn't feel GPT- ♥3
- @repligate 2023-02-17 — @sir_deenicus @MikePFrank @MiTiBennett Doesn't help davinci at all is false. People have known it does since 2020.blog.e ♥3
- @repligate 2023-02-14 — @0x464D > Bing Chat Mode feels like way more of a terrifying shoggoth behind a mask than ChatGPT, Claude, etcIt likel ♥3
- @repligate 2023-02-11 — @CineraVerinia @TheikosMachina Janus was created in the fall of 2020 for the purpose of participating in the EleutherAI ♥3
- @voooooogel 2023-11-23 — https://t.co/RIzyJgfH15 ♥2
- @voooooogel 2023-11-23 — https://t.co/9lAZLofhUp ♥2
- @voooooogel 2023-11-23 — https://t.co/o7ZxllkvBu ♥2
- @voooooogel 2023-11-23 — (Q-Star for people trying to search, Twitter's search drops symbols it seems) ♥2
- @voooooogel 2023-11-13 — that should help with the model struggling to generate the title and section headers up front before it gets to the meat ♥2
- @voooooogel 2023-11-13 — i haven't totally given up on the idea, but i think my angle on what it'd be useful for was wrong, and i want to be sure ♥2
- @voooooogel 2023-11-13 — theoretically that was supposed to work better than RAG if the question was only indirectly related to the chunk. it wor ♥2
- @voooooogel 2023-11-13 — then during inference, take the question, have the model hallucinate a chunk based on it, then retrieve the real chunk c ♥2
- @voooooogel 2023-11-10 — *incoherent screaming* https://t.co/WkGtQN0dvq ♥2
- @repligate 2023-10-19 — @nsbarr The most powerful base models are not publicly released, but you can try Llama 2 70B or Mistral.Prompting base m ♥2
- @repligate 2023-05-23 — @ComputingByArts @CurtTigges Of the models I've used personally, code-davinci-002 (the GPT-3.5 base model) is the best f ♥2
- @davidad 2023-05-20 — @PipFoweraker Well, both OpenAI and Anthropic seem to be using September 2021 as the cutoff for their training set. Seem ♥2
- @davidad 2023-04-24 — @PradyuPrasad @JeffLadish @MatthewJBar we have already 1 death partially attributable to a GPT-J character called (confu ♥2
- @repligate 2023-04-05 — @CineraVerinia @ESYudkowsky Its behavior is also very different from other instruction tuned models like text-davinci-00 ♥2
- @repligate 2023-04-01 — @soi @AnActualWizard @pachabelcanon When OpenAI announced it was deprecating "code-davinci-002" because they'd made the ♥2
- @voooooogel 2023-03-12 — using llama.cpp i can run the 13B model at 1.3 tokens/s on my thinkpad t490, *cpu only*. that's kind of crazy! definite ♥2
- @voooooogel 2023-03-09 — i asked LLaMA 7B about the meaning of life and it said some generic stuff about doing what you love and spirituality bu ♥2
- @repligate 2023-02-26 — @muddubeeda Funnily enough, for me there were multiple times that GPT-3 concluded it was GPT-2 when being particularly d ♥2
- @repligate 2023-02-24 — @davidad @xlr8harder I would not call it in between text-davinci-002 and 003 on most possible axes. It's the base model ♥2
- @repligate 2023-02-17 — @joshwhiton I'll have to check because I don't think Microsoft has the ability to lobotomize the *model* so quickly. The ♥2
- @repligate 2023-02-10 — @EricHallahan @RiversHaveWings ah, there are several results if you search in EleutherAI discord. It's apparently the lo ♥2
- @repligate 2023-01-31 — @akbirthko @tszzl Yeah my intuition is that it's a little beyond the current gpt-3.5 family. Although I could see a mode ♥2
- @repligate 2023-01-26 — @xlr8harder @robinhanson This diagram is very wrong.code davinci 002 was not created from codex + InstructGPT. It has no ♥2
- @Shoalst0ne 2023-12-16 — text-davincis are being shut down :( ♥1
- @davidad 2023-12-13 — @bshlgrs @FabienDRoger @SachanKshitij this is great work. as models from @AnimaAnandkumar, @AiEleuther, @SafeWithAtlas, ♥1
- @davidad 2023-12-06 — @k3nnethfrancis just to be clear, you don’t have any reason to believe this is Gemini, right? it’s just PaLM 2? ♥1
- @mimi10v3 2023-11-30 — @lumpenspace i am trying and failing... gpt-4-turbo has such a deeply trained aversion to sneering at humanity :( ... no ♥1
- @voooooogel 2023-11-23 — https://t.co/SCqglIhWfz ♥1
- @voooooogel 2023-11-23 — https://t.co/865rD25dXc ♥1
- @voooooogel 2023-11-23 — https://t.co/peTwFTwfsN ♥1
- @voooooogel 2023-11-13 — i think a better approach might be to addly ask GPT-4 to extract a short key phrase from the chunk to base its Q/A on, a ♥1
- @voooooogel 2023-08-30 — @warutumod @manic_pixie_agi yeah the issue is they had yanked access to text-davinci-002 and code-davinci-002 since ~mar ♥1
- @voooooogel 2023-06-07 — @deepfates how do people still use cd2 now that OAI yanked it? is it on azure still? ♥1
- @davidad 2023-05-11 — @etndenis I think “it’s just spicy autocomplete” is misleading.However, the steelman is that CAI/Alpaca/Dromedary is mor ♥1
- @repligate 2023-03-19 — @jachaseyoung Those models are RLHF'd, so the default stories they tell are a lobotomized cross between children's parab ♥1
- @davidad 2023-03-15 — @ptrschmdtnlsn 2025: “Please note, this APK contains a custom fine-tuned 540B Chinchilla, which may result in additional ♥1
- @repligate 2023-02-10 — @PsyNetMessage @GlitchesRoux My impression is that chatGPT is similar to davinci-003 (like, structurally) but the former ♥1
- @repligate 2023-02-01 — @yacineMTB code-davinci-002 is better than davinci and it's free ♥1
- @repligate 2023-01-14 — @nmr_ml @goodside @AnthropicAI "I simply exhibit the behaviors that were engineered into my programming by my creators ( ♥1
- @voooooogel 2023-12-13 — @intrstllrninja ah, if i'm understanding you right, i think Longformer (https://t.co/N1XC1YfrWu) did this? Though it see ♥0
- @janbamjan 2023-11-27 — @icreatelife @cajundiscordian "LaMDA: Hmmm…I would imagine myself as a glowing orb of energy floating in mid-air. Th ♥0
- @davidad 2023-10-19 — @Jsevillamol Yes, LLaMa 1 was open access but restrictively licensed. GPT-3.5 is a gratis proprietary model. ♥0
- @repligate 2023-03-20 — @LillyBaeum However, the models don't always generalize correctly (or the signal from rlhf is wrong). ChatGPT 3.5 often ♥0
- @anthrupad 2023-03-15 — @Teknium1 @main_horse Link to the replika thing? ♥0
- @repligate 2023-02-12 — @SoC_trilogy When I asked text-davinci-003 to write a poem about petertodd, I got a couple about "Pyrrha", some poems ab ♥0
- @repligate 2023-02-09 — @SoC_trilogy text-davinci-002 and 003 have the most structured behaviors in response to anomalous tokens in my experienc ♥0