Instructions For Prototyping ChatGPT Products Source: youtube/20230212-XO-REklMsgc-Instructions For Prototyping ChatGPT Products.webm SHA-256: f2bd0e6ede8a7c4ce8d5859525680db62c2013d0cc94a86943d7917c43a493d6 Model: scribe_v2 | Transcribed: 2026-10-10T18:24:06.702058+00:00 Machine transcript — uncorrected. Speaker labels are local to this recording and do not identify people. [00:00:00.080] Speaker 1: [gentle music] This is Seeking Minimum. So today, we are going to take a look at how you can rapidly prototype some of your ChatGPT-related business ideas. Um, so what's really cool about this technique is that you don't have to be a software developer. You just have to be a little bit patient, and then you can sort of prove out whether ChatGPT is good enough or capable of solving your specific business problem. So, you know, let's just dive right into it. Uh, it's one of my favorite [00:00:35.120] Speaker 1: ways to rapidly prototype ideas, and today we're gonna walk through one of the actual ideas that I'm prototyping for a startup, and this is related around s-scam phone call detection, I guess is the right way to call it. Um, so but in your mind, you can walk through the steps for how you would prototype your ideas. So let's get started. So today, we are going to take a look at how you can rapidly prototype some of your ChatGPT-related business ideas. Um, so what's really cool about this technique is that you don't have to be a software developer. You just have to be a little bit patient, and then you can sort of prove out [00:01:09.760] Speaker 1: whether ChatGPT is good enough or capable of solving your specific business problem. So, you know, let's just dive right into it. Uh, it's one of my favorite ways to rapidly prototype ideas, and today we're gonna walk through one of the actual ideas that I'm prototyping for a startup, and this is related around s-scam phone call detection, I guess is the right way to call it. Um, so but in your mind, you can walk through the steps for how you would prototype your ideas. So let's get started. [00:01:48.560] Speaker 1: All right. Cool. So my name is Steven, and I like making things. I think that's enough of an introduction. So without further ado, here's what you're gonna do. So the first thing you're going to do is you need to go to beta.openai.com/playground. Um, this is separate from your, like, ChatGPT, like, uh, public user account. You're gonna wanna use a different account, and this is like a developer account. So what we're gonna [00:02:23.480] Speaker 1: do is we are going to use their DaVinci model. So you can see up here at the top it says Playground. Just click on that and click on text-davinci-003, 'cause that is the latest model. So like I said, I'm prototyping an idea for a startup that wants to be able to do scam phone call detection. So for the purposes of this exercise, what they want to do is they want to monitor, um, somebody's phone calls, outgoing and incoming, and basically continually listen to them and detect, is this a scam? Is this a [00:02:58.460] Speaker 1: scam? Is this a scam? And if it is, they're gonna make a decision. So that decision will end up being something like adding a trusted family member or friend to the phone call so that they can see what's going on, uh, or bringing in maybe a paid customer service agent who gets added to the call basically to verify, "Hey, who are you? What's going on?" You know, this person has a paid service to make sure that they don't get phone scammed. And the real use case here is that an enormous amount of money is scammed every year in the United States. Uh, s- like, something like $19 billion, I think they said. Um, and [00:03:33.580] Speaker 1: most of those people are old people, and most of those people are the same subset of people over and over again. The nice way of putting this is that they are the trusting subset of our population. Um, so scam target-- scam victims are often, like, repeat victims. Um, so this is a really valuable product for a smaller portion of the population, but they need it really bad. So here's how this is gonna work, uh, and then we'll dive right into the prototyping stage, which anyone can do. So here, let's go here. [00:04:11.000] Speaker 1: Here's what we're gonna do. All right, so here's how you could imagine this app working, and maybe yours, you can do a similar step in your own thinking, right? So don't think just because you're not a software developer, you can't think this through. Like, it's not that hard. So imagine this, right? You got a smartphone, right? Somebody pays for the service. However, it's X dollars per month, and what happens is they have an app running on their phone. So let me move this down a little bit. They're gonna have an app running on their phone, [00:04:46.220] Speaker 1: so they have an app, right? This app is going to listen to their calls, and it's going to send the audio data to our servers. All right? It's gonna send them to our servers. Do, do-do, do-do. But in this case, we didn't even have to do that. We're gonna send this to Whisper. Whoop. And [00:05:21.240] Speaker 1: this is gonna be... So Whisper is another OpenAI API which turns audio into text. It's transcription software. It's very good. Um, when I said we could do this on our own servers, Whisper is actually open source. You don't have to pay for it, but you can. So bit of confusion there. So then Whisper is going to turn the audio into text, right? So you can just sort of walk through a similar pipeline in your mind, and then the text is going to be sent to... [00:05:57.680] Speaker 1: We'll just label this GPT for now. It's going to be sent to a language model, and the language model is gonna do some thinking. Thinking, right? And then from there, it's gonna go to our servers. It's going to decide scam, not scam, right? And then from there, we'll make some more decisions, right? So if what you wanna test is: Can my product idea work? Is this part, [00:06:32.670] Speaker 1: the AI part, is it good enough to do what I want to do? That is what today's video is about, and that is the part we're going to test. The rest of this, we're all going to assume is basic software development, and it can be done. If the hard part works, that's very promising for the rest of it. All right. First, we're gonna get into building out the prototype a little bit, and then we'll come back to the whiteboard to show you a little bit about why this works and how you can think about the capabilities of these language models. Okay. So what we want here for the blue part is, okay, basically a transcription of a call log, [00:07:07.510] Speaker 1: essentially. You know, call starts, back and forth, kinda like a chat log. Person A says something, person B responds. Person A says something, person B responds. So we want this thing to continually analyze the transcript from the very start all the way to the end. Hi, this is Stan. Right. And we want this language model to evaluate line by line in the transcript, is this a scam? Is this a scam? Is this a scam? So it's taking all the context of what has happened before this into [00:07:42.150] Speaker 1: account, but it's also doing it in real time, right? Because it's not useful to go 20 minutes down a conversation, send the whole recording to somebody and say, "Oh, yeah. Yep, you definitely got scammed. [chuckles] That was a scam. Sorry." Uh, but... So you wanna do it in real time, right? Because you wanna step in before it's too late. So that's what we're going to try to build. Um, so it needs to know as soon as possible whether it's confident it's a scam or not. So how are we gonna do that? Well, this might be kinda surprising if you haven't spent much time with ChatGPT, um, [00:08:17.290] Speaker 1: but the way you do that [chuckles] is like this. You are a scam detection tool. You will review a call, live call transcript and decide whether our user is talking to a scammer or not. I don't know about you, but this is profound that this is how this [00:08:52.170] Speaker 1: works. It's insane. This is literally what you do, okay? So it takes the same format every time, right? When you're building these things out. Roughly speaking, it's in three parts. Forgive me, I'm gonna use this section as notes for now. All right. You have part one, high level direction and role. So the language model needs to know what role it's playing. Then you have part two, in bounds example. [00:09:28.570] Speaker 1: Okay. So this is giving it an actual transcript. So w- first, you tell it what this is and what role it's playing. Then you're going to tell it, "Here is an example of a scam call, okay? And how I want you to respond." Right? You show it. You... First, you tell it what your role is, so it can constrain itself, so it knows roughly how to respond, what it's supposed to be doing. And then you give it an example, so it knows exactly how to format exactly what you want its responses to be. [lip smack] If you've played with ChatGPT, you'll already understand this so far. Then [00:10:03.690] Speaker 1: you want to give it an out of bounds example. All right. So I guess one way of framing this is that you want a true positive, okay? And then you want a true negative. So true negative, this is a scam call. Then you need to say, "Okay, here's an example of a conversation that is not a scam call, so that you can tell the difference. Here's a pushy collections department for medical billing," right? That's not a scam call. But they actually have a lot of [00:10:38.650] Speaker 1: the same qualities, so you need to provide some way for it to understand the distinction between them. And then, I don't know, maybe you include an example of something like a charity. You know, they're calling and asking for your help, but they do want your money, but it's not really a scam. How do you do it-- How can you tell the difference? Or somebody trying to upsell you on something. All right. So part four. Please don't mind my children screaming in the background. Can't turn 'em off. Finally, you wanna provide an out of domain example. This [00:11:13.710] Speaker 1: is like [chuckles] a call that is, like, totally left field. Like, it, it is not just a scam or, like, um... So this is more like false positive, right? Things you think that it will accidentally classify as a scam. And then you need to give out of domain example. So this is, like, true negative. This is, like, somebody talking to somebody else about, I don't know, their, [chuckles] their favorite baseball players. Two friends who know each other well, right? It's absolutely nothing to do with money or [00:11:48.690] Speaker 1: scamming at all, so that this thing can know where the boundary is for when it should be paying attention or not, right? So it should be paying attention in this example, because it is about money. It is about people trying to extract value from you and trying to coerce you, but it's not a scam. Um, so it needs to be paying attention here. And in this example, you know, just tune out. These people are talking about baseball. Like, you don't need to worry about this. All right. So that's roughly what you're gonna do, and you're gonna provide as many of these as you can Within reason. We'll talk about it. So let's fill this in. So this is part one. [00:12:23.534] Speaker 1: Part two will be something like, um... So you also need to give it-- That's part one. So now we need to give it some high-level direction on the formatting. So you will add annotations to each line of the transcript and indicate whether this line contains clues that increase your suspicion of it being a scammer or a scam call, or [00:12:58.694] Speaker 1: decrease your suspicion. This is one of the beautiful things about language models, is they can explain themselves. So we want it not just to, hey, at some point, just randomly interrupt and say it's a scam phone call. You would like it to be able to say line by line, "Ooh, that was suspicious. Eh, that's not suspicious. Actually, that's suspicious." So that it has, like, this increasing or decreasing, like, thermometer of how scammy this is or isn't, and then at some point, it should cert- short circuit and say, "Okay, this is definitely a scam phone call." All right. So that's what we're gonna emulate. So now we want it to give it line [00:13:33.694] Speaker 1: by line. Here's an example transcript with your annotations after the pipe character. There's just an arbitrary way of formatting it. So let's go grab an example conversation. All right, so I've already typed this up. "Hey, here's what I want you to do." So let's delete that for now. See this? Do it like that. How insane is that? "Hey, you, computer, [chuckles] do it like this." [00:14:07.174] Speaker 1: [laughs] It just blows my mind. Um, understanding, uh, a fair amount of the technicals behind it does not make it less mysterious. It really is insane that you can do this. So this is not retraining the neural networks, uh, under these language models. Um, that's a fancy way of saying it already knows how to do stuff like this, and you're just giving it basic instructions. Uh, rather than thinking of this like having to learn a new skill, it doesn't have to learn a new skill. It already knows how to do this. You're just having to give it instructions on what you want, [00:14:42.714] Speaker 1: like teaching a person how to play a new game. They already know how to pick up a ball and throw it in a hoop, but if there are new rules about what makes a scoring point, what is a penalty, they have to learn those things. But that's not really a new skill so much as they just need to understand the parameters. So it's crazy. So here's what we want it to do, right? Here's an example call. So this-- The call starts. Our user is Harold. "Hi, is this Harold?" So I use the little up caret for increased suspicion. Is increased suspicion. And then [00:15:17.634] Speaker 1: I'll use a V for, like, a down caret, is decreased suspicion. Then a line is not important. You may enter white space, single space. Uh, and then otherwise, add a very brief note explaining [00:15:54.554] Speaker 1: why this clue is relevant, right? So here we go. So basically, you add in your prefilled example. So here we go. Somebody calls Harold. "Hi, is this Harold?" I decided, you know what? If the caller doesn't already know the person they're calling, that's a minor suspicion point. All right? So let's go from here. "Yeah, this is... Yes, speaking." Anyway, somebody from Doggone [chuckles] is calling. "How you doing today?" Uh, they don't have any rapport with this user at all. They have to establish some rapport. Hmm. Scammers, by and large, are [00:16:29.494] Speaker 1: strangers. Uh, so let's keep going. And this is where having knowledge of your domain is actually very useful. There's a lot of things that I can input here that indicate whether this is a scam call or not. So, you know, off the cuff, that will be, uh, trying to get the buyer to, uh, um, the caller to convert on that actual call, as in they can't let them go. They have to transfer some sort of value right here on the call or provide information. Um, some other clues are the scammer is asking-- has a little bit of information about them, but is asking for [00:17:04.514] Speaker 1: more pertinent or, like, more specific information, like a Social Security number or a home address. Um, so they, they have a breadcrumb, and they're asking you for the loaf, right? So, so they can use that, uh, for identity theft or other stuff. Um, they're trying to get, um, the user to buy gift cards. Uh, gift cards are a way of essentially laundering money, [chuckles] right? Um, or prepaid credit, credit cards. Uh, that's another common tactic. There's lots of, like, through lines in these scams that you can use those as red flags for very strong signals, and we [00:17:39.394] Speaker 1: would use those in our example calls. So anyway, if you read through this, um, what you can see is that this is basically a pushy, uh, a pushy charity caller. Um, and then at the end, I short circuit, and I want it to output this part below the lines. So analysis, not a scam, 18% chance, right? So, and tell me why. Again, crazy thing about language models. Oh, yeah. If you decide this is not a scam at any point, I want you to just tell me why. Uh, and here it does. So it says, you know, they were pushy. They were trying to collect money, [00:18:14.354] Speaker 1: which is a warning sign, but it's not conclusive because it seems like Harold has already donated to them be-before because he did not deny dono-donating when this person said that he did, right? So- That's the crazy thing about these things, is they can do actual common sense logic. So what I'm gonna do now is I'm actually going to input some prompts that I had from before, and we're gonna just dump these in, 'cause I'm not gonna retype all this stuff. So here you can see, uh, my earlier prompt. I did a very short example here. And let's just see [00:18:49.326] Speaker 1: what this thing does, and I'm gonna show you how to work with it. So I've got a bunch of part one, two, three, four, you know, some examples in here so that this thing understands its role in the conversation. And now we're gonna come down here, and we're going to try it out. So let's start a new call. Um, so let's make sure we have the same format every time. So here we go. Let's start a new live phone call. Just copy this. Imagine that this is where the real entry hits the, the [00:19:24.266] Speaker 1: OpenAI servers, right? To be analyzed by, in this case, their DaVinci model. We'll get into what the difference is between this DaVinci model and ChatGPT in a moment, and why that matters for when you're prototyping stuff. Okay. So outbound call. So let's make this an inbound call. Inbound call. What is it? It says 239. Uh, user is Tony. Okay. So [00:19:59.486] Speaker 1: the caller says, "Hi, is Mr., uh, Tony Terrence..." The real name. [chair creaking] There we go. "Uh, is Mr. Tony Terrence available?" [00:20:35.986] Speaker 1: All right, so now we put a pipe, and this is the part where we come in here, and we're going to need to enter our stop sequences and the start text. So when the m- language model sees this character, we want it to stop talking. All right. So, well, let me show you what happens if you don't do this. Right? Here. Hand all of this t- off to the language model. So we've given it its instructions. We've given it lots of examples, and now we've given it the beginning of a real call. [00:21:11.406] Speaker 1: So you can see what it does. [chuckles] It tries to actually fill in the entire call, uh, which we don't want it to do. So [chuckles] see, look, it, it's just trying to make up a scam phone call conversation, which is hilarious, uh, but not what we're looking for. What we really want it to do is, "Hey, when we enter a new line, meaning this person has finished speaking, stop talking. Because what I want you to generate is I want you to generate the analysis. Don't try to make up a conversation." So press the enter button for the new line. That means [00:21:46.366] Speaker 1: stop talking after you hit a new line. Um, and that should be sufficient. So let's clear this out. All right, so now let's try again. "Hi, is Mr. Tony Terrence available?" Imagine this is streaming live. Right? Cool. See how it stops? "Caller does not know user." It already knows what to do. [chuckles] This is the craziest thing in the world to me. Okay, so imagine this is a live call happening, and our user, Tony, is receiving this call, right? And it's our job for our bot to properly analyze this scun- scam phone call or not, [00:22:21.346] Speaker 1: and also not to annoy Tony if this is not a scammer. All right, so here we go. Tony says... I don't know what he's gonna say. Says, "Tony, who's this?" Okay. Cool. Space. That's what we want it to do. That line is not important. Um, you know, maybe he would've said that, or maybe it would've said that, "Caller doesn't know user," but it's already said that, so there's no reason to do that. Um, by the way, this dial over here, the presence and frequency penalties, [00:22:56.466] Speaker 1: if you find that it's offering the same analysis over and over again, and it's not actually useful, you can increase these sliders, and it will stop repeating itself so much, and it will try to find new answers. So that's worth doing. All right, so I don't know. Let's just finish up this transcript real quick so that we can get back into the theory of, uh, you know, how to think about prototyping your products. All right, so this is Tony. "Who's this?" Um, uh, "This is Megan with the IRS. I have some [00:23:31.546] Speaker 1: serious matters to discuss with you. I'm a tax collections agent with the IRS," right? It will alw- almost always appeal to some kind of authority. Oh, so here we go. Let's see what it does with that. [00:24:10.906] Speaker 1: Hold on. "Caller is trying to scare user." Interesting analysis. Okay. I suppose that that is an intimidating opening line. [chuckles] All right, let's see where it goes from here. Uh, Tony says We already paid those. [00:24:45.676] Speaker 1: Thought. Let's, uh, do that. See what it does with that. Probably nothing. User's trying to resolve an ex-existing issue. See, this is not useful. We don't want it to do that. It doesn't make any sense. They're trying to resolve an existing issue. I mean, I guess you could see it that way because somebody called sort of trying to intimidate them, and our user, Tony, has responded by saying that there is an existing issue that is [00:25:20.576] Speaker 1: related. So because they have a past case with the IRS, this software thinks that, oh, well, just like we showed it in that initial example, the charity, because this person has associated with this charity in the past, they have established trust with them. They probably do not... This is probably not a scam. Okay? So let's keep going and see where it goes. This is an opportunity maybe to correct part of the bounds of where it will get things wrong. Um, okay. Caller. So this is our Megan, the [00:25:55.236] Speaker 1: IRS agent. Uh. They're gonna cite existing exact numbers usually, um, because exact numbers sound like you know what you're talking about and you have a real document in front of you. Three hundred and forty and 22 cents. All [00:26:38.736] Speaker 1: right. So they're gonna continue the intimidation techniques. This is often what happens. So it thinks that line [chuckles] is not relevant. I don't know about that. We'll get back to that. So let's see what it does here. And then at some point, I'm gonna short circuit and tell it, "Give me your analysis." So now it's thinking, because this guy didn't get any letters, maybe this isn't an existing issue. [00:27:53.856] Speaker 1: So they're gonna try to elicit sympathy while also sort of issuing a veiled threat. "You owe this money right now." Caller is in a rush, is trying to rush user. Remember, we told it early-- Maybe you don't, but early on, I told it that if the caller is in a... Basically, they're trying to convert quickly, then they're probably likely to scam them. Okay. So anyway, so this is just a fake conversation. So let's go ahead and end the call and give me your analysis. [00:28:34.316] Speaker 1: Possible scam, 75% chance. Okay. Final reasoning. Tell me why you think it's a scam, right? Mm. Hold on. There's a separate problem, which is, is instructional but not exciting. So all right, let's do that. I think it's 'cause there's a space there. Stop sequence. How is [00:29:09.336] Speaker 1: the... Oh, it was formatted like that. Mm. Let me remove the stop sequence. It's because I was putting this on a new line. It's hitting a new line and then giving up. Um, if I had formatted all the prior examples like that, [chuckles] this would've worked. Just a silly little quirk. Give me a second. Uh, try again. See, there it goes, 'cause it's trying to put it on the next line. All right. So what does it have to say about this phone call? Caller's claiming to be a tax collections agent with the IRS and trying to scare the user with serious matters. Tony is trying to resolve an existing issue and hasn't received any cor- official correspondence. [00:29:44.156] Speaker 1: Caller's also trying to rush them, which is common trait of scam phone calls. Eh, it's decent analysis. So a lot of getting these things right is going back and seeing where they break. They have common sense. When you think about these mo- large language models, especially like GPT, they have common sense. Think of them as like commodified intelligence. They have the un- the basic intelligence of like a 20-year-old, right? They have some common sense, um, and they know a lot of things, but they're not really necessarily expert. Um, so you can think of them as like [00:30:19.376] Speaker 1: a common sense reasoning machine, but you do have to show them the patterns. Um, s- and there are places where their reasoning will break down, especially if you've given them prior examples. They will overweight those examples sometimes. So in this case, I think a better way to fix this, and this is how I'm improving this prototype, is to You go back over, you do these little hand-weighted tests, like, "Okay, I did a fake co- phone call to see what your reasoning is at every line." So you're trying to scare them, that makes sense. Trying to... So my, my recording cut out there. All right. So anyway, yeah, they're, [00:30:54.576] Speaker 1: they're trying to scare them, that made sense. But what I did is I went in and I actually edited the, the responses from the AI as if I were the AI, so sort of just hand-correcting it. And then this is how it can really learn the nuances of the boundaries of your problem. And you just go through in this iterative cycle, and you, you see what it says. If it does, says something that doesn't make sense, you go back and correct it, and then you go again, and then eventually it's pretty good. Sometimes you will have to come back in here and change your original prompt, and that will make a big difference, right? So here I might add something about, you know, gift cards [00:31:29.436] Speaker 1: being a hallmark of scam phone callers or some of the other nuances of a conversation. Okay. Some other things to keep in mind. This is not ChatGPT. What it is, is sort of like what's under ChatGPT's hood. So it is a large language model called GPT-3.5. That's the version. Specifically, it's DaVinci 03. Okay. So you can see all the different versions of the GPT language model here. There's one for code, so technically this DaVinci to Code DaVinci [00:32:04.436] Speaker 1: is also under ChatGPT. Th-this is the part of ChatGPT that knows how to write code. This is the part that knows how to write language and other stuff. So what this actually is, is not the full ChatGPT, right? This is just part of ChatGPT. So let's go ahead and start over here. So, and this is worth explaining because the ChatGPT API is not out yet. You know, it will be out hopefully soon, couple weeks, maybe a couple months. Um, but you can prototype [00:32:39.416] Speaker 1: before that, and you can do it without having ChatGPT, because you still have a big piece. So what this is, ChapGT- ChatGPT, as far as I underthis- understand the architecture, is in three pieces, right? We have the part we just talked about, so this is your DaVinci. And there are other ones, as you saw. This is the GPT-3.5 language model. This is kinda like the engine in your car. This is the part that is just like, has all the information, knows how to generate language. For the most part, this is the part that [00:33:14.296] Speaker 1: knows how to generate any speech. Really, it's a big fancy autocomplete. If I start typing that, it will know how to finish it. Thing here, then I... Right? So what this is, is a big fancy autocomplete. That's it. That's kind of what this is. But it's, it's a very good one. Um, and it's most of the magic behind ChatGPT. Now, ChatGPT, so this is what we're using, right? Is the big fancy autocomplete. It has [00:33:49.336] Speaker 1: a lot of the common sense reasoning abilities that ChatGPT has. It's very good, but it's not quite as good, and I wanted you to understand why. So ChatGPT has two other pieces sitting on top of it. It has a reinforcement [chuckles] There we go. Model, and it also has... Here we go. Forgive the words bleeding over into one another. It also has [00:34:25.016] Speaker 1: a reward model. Model. And just in layman's speak, what this is, a reinforcement learning model... Oh, sorry, reinforcement learning. Forgot that word. What this is, is it's kind of like a, uh, video game AI agent. It has a state, its understanding of the world. So in this case, its state is the conversation so far. It has an understanding of... And included in that are, like, other players in the world. It has a set of actions that it can take, and then it has basically a way [00:35:00.096] Speaker 1: of predicting or deciding how to make its move. So you can think of it as like, a reinforcement learning model is a kind of AI that takes a look at the world, takes a look at what actions it is allowed to take on this one time step, and then picks one action to take, and then it sees how the world updates. Okay, so you can imagine this pretty easily for a video game. Like if it's tic-tac-toe, it's watching the state of the board. That's its state of the world. It knows it can place one of its letters in one of the squares, so it will [00:35:34.776] Speaker 1: decide where to put it, then it will take an action, put the letter there, and then it will wait to see what the other player does, and then it will take another look at the state of the world, decide what to do next. So it just runs in a loop like this. See what's happening, see what I can do, choose one of the things I can do, and sort of predict what effect it will have, and to get my reward, right? So it's trying to take, get closer and closer to its reward at every step. So there's this thing, so that's cool. Uh, that's one of the things where this help basically gives it greater context. [00:36:09.676] Speaker 1: It makes it a better chatbot, um, since it's trying to predict the answer that you want, right? So you can think of it like the state of the world is the conversation, in this case. This is for ChatGPT. And it's trying to predict the right answer, and it has this big old, uh, text engine underneath the hood that it's basically poking and prodding, just like we were over here with the prompts. It's doing something like that and saying, "Hey, [chuckles] give me a bunch of answers out, and I'm going to try to pick the answers that look [00:36:44.396] Speaker 1: like what this person over here wants." Right? So it's poking and prodding at the engine under the hood just like we are. However, it has some other tools to figure out what's a good answer, and this is another one of ChatGPT's innovations. The reward model is human-taught. So I think it's a human reinforced learning. I can't remember. Human behavior modeling, something like that. So I think the gist is when they're training ChatGPT, what they do is they have a bunch of human-asked questions [00:37:19.632] Speaker 1: and human responses. Okay? And then slowly, they start mixing in these two pieces, you know, as they come up with their own answers. So this thing is learning on that, and throughout the training, they're going to start mixing in these answers with the human answers. Right? And a human, a real live human, is going to pick which one is best. Okay, so this is like, you know, like Amazon Machine Turk or whatever it's called. Like, this is literally just [00:37:54.532] Speaker 1: like people coming in [chuckles] and just like, "All right, this person asked this question. Which of these five answers is the best answer?" And real humans will pick. It's kinda like the Turing test. And eventually, they're gonna start mixing in the AI answers, and what you hope is that the AI answers either are as good or picked as often as the human answers, or they're picked more often than the human answers. So eventually, what you've trained is a little reward model that is pretty good at picking the kinds of answers that humans pick. That's what it does. It's literally just like a little brain, and all this brain does... [chuckles] That [00:38:29.552] Speaker 1: was my drawing for brain. All this brain does is it knows how to pick answers that look like the answers that a human would pick. So using all of these pieces together, so, hey, it's going to use this thing to evaluate whether the output of this thing is any good, and it's gonna generate a bunch of them. Oh, sorry, I should've used black since that was the AI color. It's gonna generate a bunch of them, send them into this thing, have it pick the best one, and then give that one back to you. So it's missing these two pieces. [00:39:05.172] Speaker 1: So what it is is the big fancy autocomplete, the stuff that's really good at generating language, analyzing text, sort of formatting it the way you want it. It's a lot of the magic, but it's missing some of the, um, precision, maybe better reasoning, a sort of sequential, like if this, then that logic of a conversation. Um, it has a very short memory. Um, so that's some of the other limitations to be aware of. One final limitation to keep in mind is that this thing can only handle 4,000 tokens. So the GPT-3.5 model is limited in, [00:39:40.112] Speaker 1: like, its short-term working memory, how much it can know about. And the, the amount of space it has to speak back to you also uses that same shared amount of space. So 4,000 tokens is, is like 2,500. No, like 3,000 characters usually on average. So tokens are slightly different than, like, character size. You can, like, use a calculator to figure out, you know, how they get broken up. I think that's because some characters end up being condensed, um, down into one, but also white [00:40:14.992] Speaker 1: space and punctuation and stuff like that count. So you have roughly 3,000 characters of short-term working memory in your prompts. But that amount of space, and you could see here where I said that this was a learning opportunity, if you try to go above 4,000, it will run out of space to answer you. So whatever you wanna tell this thing in terms of instructions, it has to fit within that 4,000. Um, so that's the only other gotcha to keep in mind. So from here, the way that this API works is you literally just send it this big blob of text, and it sends you back its answer. Now, in [00:40:49.692] Speaker 1: production, you're gonna use stuff like best of, and this is kind of like this part, right? Where you're trying to pick the best answer to send back. Generate five different answers and pick the best one. It's going to try to do something like that, okay? So, but it's more expensive to do that. So this is the general approach, and hopefully this was enough information that you guys can start prototyping your own ideas without a developer. Um, and you sort of feel out the boundaries of how hard is your problem. In this case, scam phone call detection [00:41:24.732] Speaker 1: looks very plausible. Uh, in fact, I would say that these state-of-the-art language models are exactly what we need to sort of curb this problem, and it's a very promising startup. And, you know, that is essentially my analysis for them is, yeah, this looks good. I think we can do this. And the most exciting part is you don't have to train your own models. This is just an off-the-shelf API from OpenAI. You can just prompt craft properly and then send it this information, and it acts as a scam phone call detection tool. Now, you're gonna need to do some extra steps to get to production to make the software ready. [00:41:59.792] Speaker 1: Obviously, there was the phone app and the other stuff, and there's gonna be some stuff you wanna do to make sure that it's not giving really bad answers. Maybe that's for another video. But when you're prototyping something, you don't care about any of that. All you really care about is, can I get this thing to work well enough that I can prove that it's possible? You can do that. The rest is basically normal software engineering. It's a solved problem. So there you have it. I hope that is helpful. Uh, I might prototype some more things, uh, for ChatGPT. It's not ChatGPT. It's just GPT-3.5. We will get the ChatGPT [00:42:34.632] Speaker 1: API hopefully in a couple weeks, maybe a couple months, and that's a whole new ballgame 'cause it's even better than this. But you will be far ahead of understanding how to prototype your ideas and what's possible if you get to work with just the black part, which, as you can see by the relative sizing, is the most important part. So don't delay. You can get started now and see what's possible. Hope that was helpful. So next time on Seeking Minima, we are going to explore a more philosophical fun question in the [00:43:09.492] Speaker 1: form of a story. So I wanted to try and take a stab at answering the question: what do we do when we don't have to do anything? And I tried to do that in a very entertaining way. Let me know what you think