Uncorrected Scribe v2 machine transcript. Neutral speaker labels are specific to this recording. Review flags: 6.
TXT · SRT · VTT · Untouched JSON · Review flags
Source: youtube/20230212-XO-REklMsgc-Instructions For Prototyping ChatGPT Products.webm
SHA-256: f2bd0e6ede8a7c4ce8d5859525680db62c2013d0cc94a86943d7917c43a493d6
00:00:00.080 Speaker 1 [gentle music] This is Seeking Minimum.
00:00:10.460 Speaker 1 So today, we are going to take a look at how you can rapidly prototype some of your
00:00:14.740 Speaker 1 ChatGPT-related business ideas. Um, so what's really cool about this technique is
00:00:19.460 Speaker 1 that you don't have to be a software developer. You just have to be a little bit
00:00:23.220 Speaker 1 patient, and then you can sort of prove out whether ChatGPT is good enough or
00:00:28.820 Speaker 1 capable of solving your specific business problem. So, you know, let's just dive
00:00:33.600 Speaker 1 right into it. Uh, it's one of my favorite ways to rapidly prototype ideas, and
00:00:37.340 Speaker 1 today we're gonna walk through one of the actual ideas that I'm prototyping for a
00:00:40.840 Speaker 1 startup, and this is related around s-scam phone call detection, I guess is the
00:00:46.240 Speaker 1 right way to call it. Um, so but in your mind, you can walk through the steps for
00:00:51.040 Speaker 1 how you would prototype your ideas. So let's get started. So today, we are going to
00:00:55.540 Speaker 1 take a look at how you can rapidly prototype some of your ChatGPT-related business
00:01:00.300 Speaker 1 ideas. Um, so what's really cool about this technique is that you don't have to be a
00:01:04.540 Speaker 1 software developer. You just have to be a little bit patient, and then you can sort
00:01:08.860 Speaker 1 of prove out whether ChatGPT is good enough or capable of solving your specific
00:01:14.320 Speaker 1 business problem. So, you know, let's just dive right into it. Uh, it's one of my
00:01:18.640 Speaker 1 favorite ways to rapidly prototype ideas, and today we're gonna walk through one of
00:01:22.280 Speaker 1 the actual ideas that I'm prototyping for a startup, and this is related around
00:01:28.140 Speaker 1 s-scam phone call detection, I guess is the right way to call it. Um, so but in your
00:01:32.400 Speaker 1 mind, you can walk through the steps for how you would prototype your ideas. So
00:01:37.340 Speaker 1 let's get started.
00:01:48.560 Speaker 1 All right.
00:01:51.420 Speaker 1 Cool.
00:01:53.380 Speaker 1 So
00:01:55.140 Speaker 1 my name is Steven, and I like making things. I think that's enough of an
00:02:00.900 Speaker 1 introduction. So without further ado, here's what you're gonna do. So the first
00:02:06.360 Speaker 1 thing you're going to do is you need to go to beta.openai.com/playground.
00:02:12.760 Speaker 1 Um, this is separate from your, like, ChatGPT, like, uh, public user account. You're
00:02:17.560 Speaker 1 gonna wanna use a different account, and this is like a developer account. So what
00:02:23.140 Speaker 1 we're gonna do is we are going to use their DaVinci model. So
00:02:29.180 Speaker 1 you can see up here at the top it says Playground. Just click on that and click on
00:02:32.780 Speaker 1 text-davinci-003, 'cause that is the latest model.
00:02:37.420 Speaker 1 So like I said, I'm prototyping an idea for a startup that wants to be able to do
00:02:43.160 Speaker 1 scam phone call detection. So for the purposes of this exercise, what they want to
00:02:47.660 Speaker 1 do is they want to monitor, um, somebody's phone calls, outgoing and incoming, and
00:02:53.300 Speaker 1 basically continually listen to them and detect, is this a scam? Is this a scam?
00:02:59.220 Speaker 1 Is this a scam? And if it is, they're gonna make a decision. So that decision will
00:03:03.800 Speaker 1 end up being something like adding a trusted family member or friend to the phone
00:03:08.620 Speaker 1 call so that they can see what's going on, uh, or bringing in maybe a paid customer
00:03:12.760 Speaker 1 service agent who gets added to the call basically to verify, "Hey, who are you?
00:03:17.520 Speaker 1 What's going on?" You know, this person has a paid service to make sure that they
00:03:21.060 Speaker 1 don't get phone scammed. And the real use case here is that an enormous amount of
00:03:25.500 Speaker 1 money is scammed every year in the United States. Uh, s- like, something like $19
00:03:30.060 Speaker 1 billion, I think they said. Um, and most of those people are old people, and most of
00:03:35.640 Speaker 1 those people are the same subset of people over and over again. The nice way of
00:03:41.220 Speaker 1 putting this is that they are
00:03:44.140 Speaker 1 the trusting subset of our population. Um, so scam target-- scam victims
00:03:50.060 Speaker 1 are often, like, repeat victims. Um, so this is a really valuable product for a
00:03:55.780 Speaker 1 smaller portion of the population, but they need it really bad. So here's how this
00:04:00.740 Speaker 1 is gonna work, uh, and then we'll dive right into the prototyping stage, which
00:04:05.200 Speaker 1 anyone can do. So here, let's go here.
00:04:11.000 Speaker 1 Here's what we're gonna do. All right, so here's how you could imagine this app
00:04:13.580 Speaker 1 working, and maybe yours, you can do a similar step in your own thinking,
00:04:19.779 Speaker 1 right? So don't think just because you're not a software developer, you can't think
00:04:23.180 Speaker 1 this through. Like, it's not that hard. So imagine
00:04:29.040 Speaker 1 this, right? You got a smartphone, right? Somebody pays for the service. However,
00:04:32.680 Speaker 1 it's X dollars per month, and what happens is they have an app running on their
00:04:38.300 Speaker 1 phone. So
00:04:42.040 Speaker 1 let me move this down a little bit. They're gonna have an app running on their
00:04:45.300 Speaker 1 phone, so they have an app, right? This app is going to listen to their calls,
00:04:51.900 Speaker 1 and it's going to send
00:04:54.640 Speaker 1 the audio data
00:05:00.080 Speaker 1 to our servers. All right? It's gonna send them to our servers. Do,
00:05:05.980 Speaker 1 do-do, do-do.
00:05:09.780 Speaker 1 But in this case, we didn't even have to do that. We're gonna send this to
00:05:15.700 Speaker 1 Whisper.
00:05:18.900 Speaker 1 Whoop.
00:05:20.980 Speaker 1 And this is gonna be... So Whisper is another OpenAI API which turns
00:05:27.240 Speaker 1 audio into text. It's transcription software. It's very good. Um, when I said we
00:05:32.700 Speaker 1 could do this on our own servers, Whisper is actually open source. You don't have to
00:05:36.640 Speaker 1 pay for it, but you can. So bit of confusion there. So then
00:05:42.340 Speaker 1 Whisper is going to turn the audio into text,
00:05:47.440 Speaker 1 right? So you can just sort of walk through a similar pipeline in your mind, and
00:05:52.540 Speaker 1 then the text is going to be sent to...
00:05:57.680 Speaker 1 We'll just label this GPT for now. It's going to be sent to a language model, and
00:06:02.620 Speaker 1 the language model is gonna do some thinking.
00:06:10.250 Speaker 1 Thinking, right? And then from there, it's gonna go to our servers.
00:06:16.190 Speaker 1 It's going to decide scam,
00:06:20.050 Speaker 1 not scam, right? And then from there, we'll make some more decisions, right? So
00:06:26.290 Speaker 1 if what you wanna test is: Can my product idea work? Is this part,
00:06:32.670 Speaker 1 the AI part, is it good enough to do what I want to do? That is what today's video
00:06:37.770 Speaker 1 is about, and that is the part we're going to test. The rest of this, we're all
00:06:41.270 Speaker 1 going to assume is basic software development, and it can be done. If the hard part
00:06:45.550 Speaker 1 works, that's very promising for the rest of it. All right. First, we're gonna get
00:06:49.290 Speaker 1 into building out the prototype a little bit, and then we'll come back to the
00:06:52.150 Speaker 1 whiteboard to show you a little bit about why this works and how you can think about
00:06:57.030 Speaker 1 the capabilities of these language models.
00:07:00.230 Speaker 1 Okay. So what we want here for the blue part is, okay, basically a
00:07:06.070 Speaker 1 transcription of a call log, essentially. You know, call starts, back and forth,
00:07:10.610 Speaker 1 kinda like a chat log. Person A says something, person B responds. Person A says
00:07:14.669 Speaker 1 something, person B responds. So we want this thing to continually analyze the
00:07:19.930 Speaker 1 transcript from the very start all the way to the end.
00:07:25.970 Speaker 1 Hi, this is Stan. Right. And we want this language model to evaluate
00:07:32.030 Speaker 1 line by line in the transcript, is this a scam? Is this a scam? Is this a scam?
00:07:38.290 Speaker 1 So it's taking all the context of what has happened before this into account, but
00:07:43.250 Speaker 1 it's also doing it in real time, right? Because it's not useful to go 20 minutes
00:07:47.930 Speaker 1 down a conversation, send the whole recording to somebody and say, "Oh, yeah. Yep,
00:07:53.430 Speaker 1 you definitely got scammed. [chuckles] That was a scam. Sorry." Uh, but... So you
00:07:58.970 Speaker 1 wanna do it in real time, right? Because you wanna step in before it's too late. So
00:08:03.910 Speaker 1 that's what we're going to try to build. Um, so it needs to know as soon as possible
00:08:07.950 Speaker 1 whether it's confident it's a scam or not. So how are we gonna do that?
00:08:12.610 Speaker 1 Well, this might be kinda surprising if you haven't spent much time with ChatGPT,
00:08:16.810 Speaker 1 um, but the way you do that [chuckles] is like this. You
00:08:22.670 Speaker 1 are a scam detection tool.
00:08:29.730 Speaker 1 You will review a call, live call
00:08:36.150 Speaker 1 transcript and decide whether
00:08:42.050 Speaker 1 our user is talking to a scammer or not.
00:08:49.410 Speaker 1 I don't know about you, but this is profound that this is how this works. It's
00:08:52.690 Speaker 1 insane. This is literally what you do, okay? So it takes the same format
00:08:58.690 Speaker 1 every time, right? When you're building these things out. Roughly speaking, it's in
00:09:02.370 Speaker 1 three parts. Forgive me, I'm gonna use this section as notes for now. All right. You
00:09:06.570 Speaker 1 have part one,
00:09:09.330 Speaker 1 high level direction and role. So the language model needs
00:09:15.230 Speaker 1 to know what role it's playing. Then you have part two,
00:09:22.190 Speaker 1 in
00:09:24.790 Speaker 1 bounds example.
00:09:28.570 Speaker 1 Okay. So this is giving it an actual transcript. So w- first,
00:09:34.830 Speaker 1 you tell it what this is and what role it's playing. Then you're going to tell it,
00:09:39.470 Speaker 1 "Here is an example of a scam call, okay? And how I want you to
00:09:45.050 Speaker 1 respond." Right? You show it. You... First, you tell it what your role is, so it can
00:09:48.750 Speaker 1 constrain itself, so it knows roughly how to respond, what it's supposed to be
00:09:53.310 Speaker 1 doing. And then you give it an example, so it knows exactly how to format exactly
00:09:57.630 Speaker 1 what you want its responses to be. [lip smack] If you've played with ChatGPT, you'll
00:10:01.210 Speaker 1 already understand this so far. Then you want to give it
00:10:06.590 Speaker 1 an out of bounds example.
00:10:12.750 Speaker 1 All right. So I guess one way of framing this is that you want a true positive,
00:10:20.290 Speaker 1 okay? And then you want a true negative. So true negative, this is a scam call. Then
00:10:25.810 Speaker 1 you need to say, "Okay, here's an example of a conversation that is not a scam call,
00:10:30.650 Speaker 1 so that you can tell the difference. Here's a pushy collections department for
00:10:35.090 Speaker 1 medical billing," right? That's not a scam call. But they actually have a lot of the
00:10:38.770 Speaker 1 same qualities, so you need to provide some way for it to understand the distinction
00:10:42.890 Speaker 1 between them. And then, I don't know, maybe you include an example of something like
00:10:47.610 Speaker 1 a charity. You know, they're calling and asking for your help, but they do want your
00:10:50.830 Speaker 1 money, but it's not really a scam. How do you do it-- How can you tell the
00:10:54.790 Speaker 1 difference? Or somebody trying to upsell you on something.
00:10:59.250 Speaker 1 All right. So part four.
00:11:05.150 Speaker 1 Please don't mind my children screaming in the background. Can't turn 'em off.
00:11:09.490 Speaker 1 Finally, you wanna provide an out of domain example. This is like
00:11:14.310 Speaker 1 [chuckles] a call that is, like, totally left field. Like, it, it is not just a scam
00:11:20.510 Speaker 1 or, like, um... So this is more like false
00:11:26.270 Speaker 1 positive, right? Things you think that it will accidentally classify as a scam.
00:11:32.410 Speaker 1 And then you need to give out of domain example. So this is, like, true negative.
00:11:37.770 Speaker 1 This is, like, somebody talking to somebody else about, I don't know, their,
00:11:42.730 Speaker 1 [chuckles] their favorite baseball players. Two friends who know each other well,
00:11:45.810 Speaker 1 right? It's absolutely nothing to do with money or scamming at all, so that this
00:11:50.630 Speaker 1 thing can know where the boundary is for when it should be paying attention or not,
00:11:55.350 Speaker 1 right? So it should be paying attention in this example, because it is about money.
00:11:59.890 Speaker 1 It is about people trying to extract value from you and trying to coerce you, but
00:12:04.330 Speaker 1 it's not a scam. Um, so it needs to be paying attention here. And in this example,
00:12:10.210 Speaker 1 you know, just tune out. These people are talking about baseball. Like, you don't
00:12:13.210 Speaker 1 need to worry about this. All right. So that's roughly what you're gonna do, and
00:12:16.570 Speaker 1 you're gonna provide as many of these as you can Within reason. We'll talk about it.
00:12:21.634 Speaker 1 So let's fill this in. So this is part one. Part two will be something like, um...
00:12:27.194 Speaker 1 So you also need to give it--
00:12:30.894 Speaker 1 That's part one. So now we need to give it some high-level direction on the
00:12:35.134 Speaker 1 formatting. So you will add annotations to each line of the
00:12:40.874 Speaker 1 transcript
00:12:43.134 Speaker 1 and indicate whether this line contains clues
00:12:49.314 Speaker 1 that increase your suspicion
00:12:54.974 Speaker 1 of it being a scammer or a scam call, or decrease your suspicion.
00:13:01.594 Speaker 1 This is one of the beautiful things about language models, is they can explain
00:13:04.174 Speaker 1 themselves. So we want it not just to, hey, at some point, just randomly interrupt
00:13:09.774 Speaker 1 and say it's a scam phone call. You would like it to be able to say line by line,
00:13:13.574 Speaker 1 "Ooh, that was suspicious. Eh, that's not suspicious. Actually, that's suspicious."
00:13:18.974 Speaker 1 So that it has, like, this increasing or decreasing, like, thermometer of how scammy
00:13:23.514 Speaker 1 this is or isn't, and then at some point, it should cert- short circuit and say,
00:13:27.234 Speaker 1 "Okay, this is definitely a scam phone call." All right. So that's what we're gonna
00:13:30.694 Speaker 1 emulate. So now we want it to give it line by line. Here's an example
00:13:38.014 Speaker 1 transcript with your annotations
00:13:43.994 Speaker 1 after the pipe character. There's just an arbitrary way of formatting it. So let's
00:13:49.374 Speaker 1 go grab an example conversation. All right, so I've already typed this up. "Hey,
00:13:55.514 Speaker 1 here's what I want you to do." So let's delete that for now.
00:14:00.094 Speaker 1 See this? Do it like that. How insane is that? "Hey, you, computer,
00:14:05.834 Speaker 1 [chuckles] do it like this." [laughs] It just blows my mind. Um,
00:14:11.974 Speaker 1 understanding, uh, a fair amount of the technicals behind it does not make it less
00:14:15.574 Speaker 1 mysterious. It really is insane that you can do this. So this is not retraining the
00:14:21.314 Speaker 1 neural networks, uh, under these language models. Um, that's a fancy way of saying
00:14:26.594 Speaker 1 it already knows how to do stuff like this, and you're just giving it basic
00:14:29.894 Speaker 1 instructions. Uh, rather than thinking of this like
00:14:35.734 Speaker 1 having to learn a new skill, it doesn't have to learn a new skill. It already knows
00:14:39.114 Speaker 1 how to do this. You're just having to give it instructions on what you want, like
00:14:42.954 Speaker 1 teaching a person how to play a new game. They already know how to pick up a ball
00:14:46.194 Speaker 1 and throw it in a hoop, but if there are new rules about what makes a scoring point,
00:14:50.834 Speaker 1 what is a penalty, they have to learn those things. But that's not really a new
00:14:54.694 Speaker 1 skill so much as they just need to understand the parameters. So it's crazy. So
00:15:00.274 Speaker 1 here's what we want it to do, right? Here's an example call. So this-- The call
00:15:04.094 Speaker 1 starts. Our user is Harold. "Hi, is this Harold?" So I use the little up caret for
00:15:10.054 Speaker 1 increased suspicion.
00:15:12.754 Speaker 1 Is
00:15:14.614 Speaker 1 increased suspicion. And then I'll use a V for, like, a down caret, is decreased
00:15:21.274 Speaker 1 suspicion.
00:15:27.054 Speaker 1 Then
00:15:30.974 Speaker 1 a line is not important.
00:15:34.854 Speaker 1 You may enter white space,
00:15:39.314 Speaker 1 single space.
00:15:42.114 Speaker 1 Uh, and then otherwise, add a very brief
00:15:48.514 Speaker 1 note explaining
00:15:54.554 Speaker 1 why this clue is relevant, right?
00:15:58.594 Speaker 1 So here we go. So basically, you add in your prefilled example. So here we go.
00:16:03.814 Speaker 1 Somebody calls Harold. "Hi, is this Harold?" I decided, you know what? If the caller
00:16:08.954 Speaker 1 doesn't already know the person they're calling,
00:16:12.374 Speaker 1 that's a minor suspicion point. All right? So let's go from here. "Yeah, this is...
00:16:18.074 Speaker 1 Yes, speaking." Anyway, somebody from Doggone [chuckles] is calling. "How you doing
00:16:23.354 Speaker 1 today?" Uh, they don't have any rapport with this user at all. They have to
00:16:26.294 Speaker 1 establish some rapport. Hmm. Scammers, by and large, are strangers. Uh, so let's
00:16:32.054 Speaker 1 keep going. And this is where having knowledge of your domain is actually very
00:16:35.634 Speaker 1 useful. There's a lot of things that I can input here that indicate whether this is
00:16:40.434 Speaker 1 a scam call or not. So, you know, off the cuff, that will be, uh, trying to get the
00:16:45.674 Speaker 1 buyer to, uh, um, the caller to convert on that actual call, as in they can't
00:16:51.654 Speaker 1 let them go. They have to transfer some sort of value right here on the call or
00:16:55.574 Speaker 1 provide information. Um, some other clues are the scammer is asking--
00:17:01.654 Speaker 1 has a little bit of information about them, but is asking for more pertinent or,
00:17:06.114 Speaker 1 like, more specific information, like a Social Security number or a home address.
00:17:11.194 Speaker 1 Um, so they, they have a breadcrumb, and they're asking you for the loaf, right? So,
00:17:16.474 Speaker 1 so they can use that, uh, for identity theft or other stuff. Um, they're trying to
00:17:20.774 Speaker 1 get, um, the user to buy gift cards. Uh, gift cards are a way of essentially
00:17:26.754 Speaker 1 laundering money, [chuckles] right? Um, or prepaid credit, credit cards. Uh, that's
00:17:31.854 Speaker 1 another common tactic. There's lots of, like, through lines in these scams that you
00:17:36.374 Speaker 1 can use those as red flags for very strong signals, and we would use those in our
00:17:40.354 Speaker 1 example calls. So anyway, if you read through this, um, what you can see is that
00:17:45.314 Speaker 1 this is basically a pushy,
00:17:48.654 Speaker 1 uh, a pushy charity caller. Um, and then at the end, I short circuit, and I want
00:17:54.594 Speaker 1 it to output this part below the lines. So analysis, not a scam, 18% chance, right?
00:18:00.794 Speaker 1 So, and tell me why. Again, crazy thing about language models. Oh, yeah. If you
00:18:04.834 Speaker 1 decide this is not a scam at any point, I want you to just tell me why. Uh, and here
00:18:09.574 Speaker 1 it does. So it says, you know, they were pushy. They were trying to collect money,
00:18:14.354 Speaker 1 which is a warning sign, but it's not conclusive because it seems like Harold has
00:18:17.814 Speaker 1 already donated to them be-before because he did not deny dono-donating when this
00:18:22.394 Speaker 1 person said that he did, right? So- That's the crazy thing about these things, is
00:18:27.646 Speaker 1 they can do actual common sense logic. So what I'm gonna do now
00:18:33.646 Speaker 1 is I'm actually going to input some prompts that I had from before, and we're gonna
00:18:38.426 Speaker 1 just dump these in, 'cause I'm not gonna retype all this stuff. So here you can see,
00:18:43.966 Speaker 1 uh, my earlier prompt. I did a very short example here. And let's just see what this
00:18:49.566 Speaker 1 thing does, and I'm gonna show you how to work with it. So I've got a bunch of part
00:18:53.706 Speaker 1 one, two, three, four, you know, some examples in here so that this thing
00:18:57.666 Speaker 1 understands its role in the conversation. And now we're gonna come down here, and
00:19:02.486 Speaker 1 we're going to try it out. So let's start a new call. Um, so
00:19:08.406 Speaker 1 let's make sure we have the same format every time. So here we go. Let's start a new
00:19:12.706 Speaker 1 live phone call.
00:19:15.586 Speaker 1 Just copy this.
00:19:19.386 Speaker 1 Imagine that this is where the real entry hits the, the OpenAI servers, right?
00:19:26.086 Speaker 1 To be analyzed by, in this case, their DaVinci model. We'll get into what the
00:19:30.106 Speaker 1 difference is between this DaVinci model and ChatGPT in a moment, and why that
00:19:34.306 Speaker 1 matters for when you're prototyping stuff.
00:19:38.266 Speaker 1 Okay. So
00:19:43.186 Speaker 1 outbound call. So let's make this an inbound call.
00:19:47.266 Speaker 1 Inbound call.
00:19:50.166 Speaker 1 What is it? It says 239.
00:19:53.706 Speaker 1 Uh, user is Tony.
00:19:56.686 Speaker 1 Okay.
00:19:58.806 Speaker 1 So the caller says,
00:20:01.126 Speaker 1 "Hi,
00:20:06.866 Speaker 1 is Mr.,
00:20:12.306 Speaker 1 uh,
00:20:14.346 Speaker 1 Tony Terrence..."
00:20:16.766 Speaker 1 The real name.
00:20:17.506 Speaker 1 [chair creaking]
00:20:31.326 Speaker 1 There we go. "Uh, is Mr. Tony Terrence available?"
00:20:35.986 Speaker 1 All right, so now we put a pipe, and this is the part where
00:20:40.766 Speaker 1 we come in here, and we're going to need to enter our stop sequences and the start
00:20:45.446 Speaker 1 text. So when the m- language model sees this character,
00:20:51.746 Speaker 1 we want it to stop
00:20:55.346 Speaker 1 talking. All right. So, well, let me show you what happens if you don't do this.
00:20:58.846 Speaker 1 Right? Here. Hand all of this t- off to the language model. So we've given it its
00:21:03.166 Speaker 1 instructions. We've given it lots of examples, and now we've given it the beginning
00:21:06.506 Speaker 1 of a real call.
00:21:11.406 Speaker 1 So you can see what it does. [chuckles] It tries to actually fill in
00:21:17.426 Speaker 1 the entire call, uh, which we don't want it to do. So [chuckles] see, look, it,
00:21:23.526 Speaker 1 it's just trying to make up a scam phone call conversation, which is hilarious, uh,
00:21:27.906 Speaker 1 but not what we're looking for. What we really want it to do is, "Hey, when we enter
00:21:31.546 Speaker 1 a new line, meaning this person has finished speaking, stop talking. Because what I
00:21:36.566 Speaker 1 want you to generate is I want you to generate the analysis. Don't try to make up a
00:21:41.946 Speaker 1 conversation." So press the enter button for the new line. That means stop talking
00:21:47.266 Speaker 1 after you hit a new line. Um, and that should be sufficient. So let's clear this
00:21:52.046 Speaker 1 out.
00:21:54.106 Speaker 1 All right, so now let's try again. "Hi, is Mr. Tony Terrence available?" Imagine
00:21:58.266 Speaker 1 this is streaming live. Right?
00:22:01.786 Speaker 1 Cool. See how it stops? "Caller does not know user." It already knows what to do.
00:22:07.606 Speaker 1 [chuckles] This is the craziest thing in the world to me. Okay, so imagine this is a
00:22:11.386 Speaker 1 live call happening, and our user, Tony, is receiving this call, right? And it's our
00:22:15.946 Speaker 1 job for our bot to properly analyze this scun- scam phone call or not, and also
00:22:21.886 Speaker 1 not to annoy Tony if this is not a scammer. All right, so here we go. Tony says... I
00:22:27.926 Speaker 1 don't know what he's gonna say.
00:22:30.806 Speaker 1 Says, "Tony, who's this?"
00:22:35.066 Speaker 1 Okay.
00:22:41.486 Speaker 1 Cool. Space. That's what we want it to do. That line is not important. Um, you know,
00:22:47.286 Speaker 1 maybe he would've said that, or maybe it would've said that, "Caller doesn't know
00:22:50.026 Speaker 1 user," but it's already said that, so there's no reason to do that. Um, by the way,
00:22:53.626 Speaker 1 this dial over here, the presence and frequency penalties, if you find that it's
00:22:57.426 Speaker 1 offering the same analysis over and over again, and it's not actually useful, you
00:23:00.866 Speaker 1 can increase these sliders, and it will stop repeating itself so much, and it will
00:23:06.446 Speaker 1 try to find new answers. So that's worth doing.
00:23:10.586 Speaker 1 All right, so I don't know. Let's just finish up this transcript real quick so that
00:23:13.566 Speaker 1 we can get back into the theory of, uh, you know, how to think about prototyping
00:23:18.286 Speaker 1 your products. All right, so this is Tony. "Who's this?" Um,
00:23:23.986 Speaker 1 uh, "This is Megan with the IRS.
00:23:29.766 Speaker 1 I have some serious matters to discuss with you.
00:23:44.646 Speaker 1 I'm a tax collections agent
00:23:50.846 Speaker 1 with the IRS," right?
00:23:54.086 Speaker 1 It will alw- almost always appeal to some kind of authority.
00:24:01.546 Speaker 1 Oh, so here we go. Let's see what it does with that.
00:24:10.906 Speaker 1 Hold on. "Caller
00:24:16.926 Speaker 1 is trying to scare user." Interesting analysis. Okay. I suppose that that
00:24:22.946 Speaker 1 is an intimidating opening line. [chuckles] All right, let's see where it goes from
00:24:26.506 Speaker 1 here. Uh, Tony says
00:24:39.216 Speaker 1 We already paid those.
00:24:45.676 Speaker 1 Thought.
00:24:47.616 Speaker 1 Let's, uh,
00:24:49.776 Speaker 1 do that. See what it does with that.
00:24:53.156 Speaker 1 Probably nothing.
00:24:55.936 Speaker 1 User's trying to resolve an ex-existing issue. See, this is not useful. We
00:25:01.816 Speaker 1 don't want it to do that. It doesn't make any sense. They're trying to resolve an
00:25:06.196 Speaker 1 existing issue. I mean, I guess you could see it that way because somebody
00:25:11.936 Speaker 1 called sort of trying to intimidate them, and our user, Tony, has responded
00:25:17.876 Speaker 1 by saying that there is an existing issue that is related. So because they have a
00:25:22.376 Speaker 1 past case with the IRS, this software thinks that, oh, well, just like we
00:25:28.416 Speaker 1 showed it in that initial example,
00:25:31.876 Speaker 1 the charity, because this person has associated with this charity in the past, they
00:25:35.696 Speaker 1 have established trust with them. They probably do not... This is probably not a
00:25:41.055 Speaker 1 scam. Okay? So let's keep going and see where it goes.
00:25:45.676 Speaker 1 This is an opportunity maybe to correct part of the bounds of where it will get
00:25:49.536 Speaker 1 things wrong.
00:25:51.856 Speaker 1 Um, okay. Caller. So this is our Megan, the IRS agent. Uh.
00:26:04.896 Speaker 1 They're gonna cite existing exact numbers usually, um, because exact numbers sound
00:26:10.076 Speaker 1 like you know what you're talking about and you have a real document in front of
00:26:12.316 Speaker 1 you.
00:26:15.116 Speaker 1 Three hundred and forty and 22 cents.
00:26:24.096 Speaker 1 All
00:26:38.736 Speaker 1 right. So they're gonna continue the intimidation techniques. This is often what
00:26:42.716 Speaker 1 happens.
00:26:45.336 Speaker 1 So it thinks that line [chuckles] is not relevant. I don't know about that. We'll
00:26:49.196 Speaker 1 get back to that.
00:26:55.096 Speaker 1 So let's see what it does here.
00:26:59.536 Speaker 1 And then at some point, I'm gonna short circuit and tell it, "Give me your
00:27:01.676 Speaker 1 analysis."
00:27:06.676 Speaker 1 So now it's thinking, because this guy didn't get any letters, maybe this isn't an
00:27:11.396 Speaker 1 existing issue.
00:27:53.856 Speaker 1 So they're gonna try to elicit sympathy while also sort of issuing a veiled threat.
00:27:57.236 Speaker 1 "You owe this money right now."
00:28:04.096 Speaker 1 Caller is in a rush, is trying to rush user. Remember, we told it early-- Maybe you
00:28:09.316 Speaker 1 don't, but early on, I told it that if the caller is in a... Basically, they're
00:28:14.636 Speaker 1 trying to convert quickly, then they're probably likely to scam them. Okay.
00:28:20.676 Speaker 1 So anyway, so this is just a fake conversation. So let's go ahead and end the call
00:28:26.136 Speaker 1 and give me your analysis.
00:28:34.316 Speaker 1 Possible scam, 75% chance. Okay. Final reasoning. Tell me why you
00:28:40.276 Speaker 1 think it's a scam, right?
00:28:47.736 Speaker 1 Mm. Hold on. There's a separate problem, which is, is instructional but not
00:28:53.416 Speaker 1 exciting. So
00:28:56.555 Speaker 1 all right, let's do that.
00:29:02.396 Speaker 1 I think it's 'cause there's a space there.
00:29:08.456 Speaker 1 Stop sequence. How is the... Oh, it was formatted like that. Mm. Let me remove the
00:29:13.076 Speaker 1 stop sequence. It's because I was putting this on a new line. It's hitting a new
00:29:17.696 Speaker 1 line and then giving up. Um, if I had formatted all the prior examples like that,
00:29:22.316 Speaker 1 [chuckles] this would've worked. Just a silly little quirk. Give me a second. Uh,
00:29:27.276 Speaker 1 try again.
00:29:30.456 Speaker 1 See, there it goes, 'cause it's trying to put it on the next line. All right. So
00:29:33.136 Speaker 1 what does it have to say about this phone call? Caller's claiming to be a tax
00:29:36.096 Speaker 1 collections agent with the IRS and trying to scare the user with serious matters.
00:29:39.336 Speaker 1 Tony is trying to resolve an existing issue and hasn't received any cor- official
00:29:42.456 Speaker 1 correspondence. Caller's also trying to rush them, which is common trait of scam
00:29:46.696 Speaker 1 phone calls. Eh, it's decent analysis. So a lot of getting these things right is
00:29:51.936 Speaker 1 going back and seeing where they break. They have common sense. When you think about
00:29:56.676 Speaker 1 these mo- large language models, especially like GPT,
00:30:00.896 Speaker 1 they have common sense. Think of them as like commodified intelligence. They have
00:30:04.316 Speaker 1 the un- the basic intelligence of like a 20-year-old, right? They have some common
00:30:09.976 Speaker 1 sense, um, and they know a lot of things, but they're not really
00:30:15.716 Speaker 1 necessarily expert. Um, so you can think of them as like a common sense reasoning
00:30:20.376 Speaker 1 machine, but you do have to show them the patterns. Um, s- and there are places
00:30:24.976 Speaker 1 where their reasoning will break down, especially if you've given them prior
00:30:27.996 Speaker 1 examples. They will overweight those examples sometimes. So in this case, I think a
00:30:33.656 Speaker 1 better way to fix this, and this is how I'm improving this prototype, is to You go
00:30:39.316 Speaker 1 back over, you do these little hand-weighted tests, like, "Okay, I did a fake co-
00:30:43.176 Speaker 1 phone call to see what your reasoning is at every line." So you're trying to scare
00:30:46.815 Speaker 1 them, that makes sense. Trying to... So my, my recording cut out there. All right.
00:30:52.476 Speaker 1 So anyway, yeah, they're, they're trying to scare them, that made sense. But what I
00:30:56.476 Speaker 1 did is I went in and I actually edited the, the responses from the AI as if I were
00:31:02.356 Speaker 1 the AI, so sort of just hand-correcting it. And then this is how it can really learn
00:31:06.896 Speaker 1 the nuances of the boundaries of your problem. And you just go through in this
00:31:10.316 Speaker 1 iterative cycle, and you, you see what it says. If it does, says something that
00:31:14.136 Speaker 1 doesn't make sense, you go back and correct it, and then you go again, and then
00:31:17.316 Speaker 1 eventually it's pretty good. Sometimes you will have to come back in here and change
00:31:21.916 Speaker 1 your original prompt, and that will make a big difference, right? So here I might
00:31:27.576 Speaker 1 add something about, you know, gift cards being a hallmark of scam phone callers or
00:31:31.876 Speaker 1 some of the other nuances of a conversation.
00:31:35.596 Speaker 1 Okay. Some other things to keep in mind.
00:31:38.896 Speaker 1 This is not ChatGPT. What it is, is sort of like what's under ChatGPT's
00:31:44.916 Speaker 1 hood. So it is a large language model called GPT-3.5.
00:31:50.976 Speaker 1 That's the version. Specifically, it's DaVinci 03. Okay. So you can see all the
00:31:56.336 Speaker 1 different versions of the GPT language model here. There's one for code, so
00:32:01.676 Speaker 1 technically this DaVinci to Code DaVinci is also under ChatGPT. Th-this is the part
00:32:06.516 Speaker 1 of ChatGPT that knows how to write code. This is the part that knows how to write
00:32:10.096 Speaker 1 language and other stuff. So what this actually is, is not the full
00:32:16.336 Speaker 1 ChatGPT, right? This is just part of ChatGPT.
00:32:22.216 Speaker 1 So let's go ahead and start over here.
00:32:28.876 Speaker 1 So, and this is worth explaining because the ChatGPT API is not out yet. You know,
00:32:34.776 Speaker 1 it will be out hopefully soon, couple weeks, maybe a couple months. Um, but you can
00:32:38.856 Speaker 1 prototype before that, and you can do it without having ChatGPT, because you still
00:32:44.576 Speaker 1 have a big piece. So what this is, ChapGT- ChatGPT, as far as I underthis-
00:32:48.776 Speaker 1 understand the architecture, is in three pieces, right? We have the part we just
00:32:53.476 Speaker 1 talked about, so this is your DaVinci.
00:32:57.916 Speaker 1 And there are other ones, as you saw. This is the GPT-3.5 language model. This is
00:33:03.476 Speaker 1 kinda like the engine in your car. This is the part that is just like, has all
00:33:09.436 Speaker 1 the information, knows how to generate language. For the most part, this is the part
00:33:13.976 Speaker 1 that knows how to generate any speech. Really, it's a big fancy autocomplete.
00:33:20.176 Speaker 1 If I start typing that, it will know
00:33:26.736 Speaker 1 how to finish it.
00:33:29.276 Speaker 1 Thing here, then I... Right? So what this is, is a big fancy autocomplete.
00:33:35.676 Speaker 1 That's it. That's kind of what this is. But it's, it's a very good one. Um,
00:33:41.896 Speaker 1 and it's most of the magic behind ChatGPT. Now, ChatGPT, so this is what we're
00:33:46.536 Speaker 1 using, right? Is the big fancy autocomplete. It has a lot of the common sense
00:33:50.516 Speaker 1 reasoning abilities that ChatGPT has. It's very good, but it's not quite as good,
00:33:56.636 Speaker 1 and I wanted you to understand why. So ChatGPT
00:34:02.536 Speaker 1 has two other pieces sitting on top of it. It has a
00:34:08.836 Speaker 1 reinforcement [chuckles] There we go.
00:34:15.296 Speaker 1 Model, and it also has... Here we go.
00:34:20.536 Speaker 1 Forgive the words bleeding over into one another. It also has a reward model.
00:34:28.656 Speaker 1 Model. And just in layman's speak, what this is, a reinforcement learning model...
00:34:34.476 Speaker 1 Oh, sorry, reinforcement learning. Forgot that word. What this is, is it's kind of
00:34:39.416 Speaker 1 like a, uh, video game AI agent. It has a state, its understanding of
00:34:45.416 Speaker 1 the world. So in this case, its state is the conversation so far. It has an
00:34:49.716 Speaker 1 understanding of... And included in that are, like, other players in the world. It
00:34:54.795 Speaker 1 has a set of actions that it can take, and then it has basically a way of predicting
00:35:01.136 Speaker 1 or deciding how to make its move. So you can think of it as like, a reinforcement
00:35:06.676 Speaker 1 learning model is a kind of AI that takes a look at the world,
00:35:11.496 Speaker 1 takes a look at what actions it is allowed to take on this one time step, and then
00:35:16.156 Speaker 1 picks one action to take, and then it sees how the world updates.
00:35:22.156 Speaker 1 Okay, so you can imagine this pretty easily for a video game. Like if it's
00:35:25.436 Speaker 1 tic-tac-toe, it's watching the state of the board. That's its state of the world. It
00:35:29.276 Speaker 1 knows it can place one of its letters in one of the squares, so it will decide
00:35:35.296 Speaker 1 where to put it, then it will take an action, put the letter there, and then it will
00:35:38.976 Speaker 1 wait to see what the other player does, and then it will take another look at the
00:35:42.316 Speaker 1 state of the world, decide what to do next. So it just runs in a loop like this. See
00:35:46.736 Speaker 1 what's happening, see what I can do, choose one of the things I can do, and sort of
00:35:51.896 Speaker 1 predict what effect it will have, and to get my reward, right? So it's trying to
00:35:56.496 Speaker 1 take, get closer and closer to its reward at every step. So there's this thing,
00:36:02.816 Speaker 1 so that's cool. Uh, that's one of the things where this help basically gives it
00:36:08.476 Speaker 1 greater context. It makes it a better chatbot, um, since it's trying to
00:36:14.156 Speaker 1 predict the answer that you want, right? So you can think of it like the state of
00:36:20.056 Speaker 1 the world is the conversation, in this case. This is for ChatGPT. And it's trying to
00:36:24.956 Speaker 1 predict the right answer, and it has this big old, uh, text engine
00:36:30.936 Speaker 1 underneath the hood that it's basically poking and prodding, just like we were over
00:36:35.416 Speaker 1 here with the prompts. It's doing something like that and saying, "Hey, [chuckles]
00:36:39.876 Speaker 1 give me a bunch of answers out, and I'm going to try to pick the answers that look
00:36:44.396 Speaker 1 like what this person over here wants." Right? So it's poking and prodding at
00:36:50.352 Speaker 1 the engine under the hood just like we are. However, it has some other tools to
00:36:55.252 Speaker 1 figure out what's a good answer, and this is another one of ChatGPT's innovations.
00:36:59.932 Speaker 1 The reward model is human-taught. So I think it's a
00:37:05.572 Speaker 1 human reinforced learning. I can't remember. Human behavior modeling, something like
00:37:10.332 Speaker 1 that. So I think the gist is when they're training ChatGPT, what they do is
00:37:16.632 Speaker 1 they have a bunch of human-asked questions and human responses. Okay?
00:37:22.732 Speaker 1 And then slowly, they start mixing in these two pieces, you know, as they come up
00:37:28.372 Speaker 1 with their own answers. So this thing is learning on that, and throughout the
00:37:32.472 Speaker 1 training, they're going to start mixing in these answers with the human answers.
00:37:38.632 Speaker 1 Right?
00:37:43.092 Speaker 1 And a human, a real live human, is going to pick which one is best. Okay,
00:37:49.012 Speaker 1 so this is like, you know, like Amazon Machine Turk or whatever it's called. Like,
00:37:53.572 Speaker 1 this is literally just like people coming in [chuckles] and just like, "All right,
00:37:57.592 Speaker 1 this person asked this question. Which of these five answers is the best answer?"
00:38:01.572 Speaker 1 And real humans will pick. It's kinda like the Turing test. And eventually, they're
00:38:05.512 Speaker 1 gonna start mixing in the AI answers, and what you hope is that the AI answers
00:38:10.332 Speaker 1 either are as good or picked as often as the human answers, or they're picked more
00:38:15.252 Speaker 1 often than the human answers. So eventually, what you've trained is a little reward
00:38:19.512 Speaker 1 model that is pretty good at picking the kinds of answers that humans pick. That's
00:38:24.652 Speaker 1 what it does. It's literally just like a little brain, and all this brain does...
00:38:28.992 Speaker 1 [chuckles] That was my drawing for brain. All this brain does is it knows how to
00:38:34.132 Speaker 1 pick answers that look like the answers that a human would pick. So using all of
00:38:38.632 Speaker 1 these pieces together, so, hey, it's going to use this thing to evaluate whether the
00:38:44.552 Speaker 1 output of this thing is any good, and it's gonna generate a bunch of them. Oh,
00:38:49.252 Speaker 1 sorry, I should've used black since that was the AI color. It's gonna generate a
00:38:52.712 Speaker 1 bunch of them,
00:38:55.412 Speaker 1 send them into this thing, have it pick the best one, and then give that one back to
00:39:00.032 Speaker 1 you. So it's missing these two pieces. So what it is is the
00:39:05.972 Speaker 1 big fancy autocomplete, the stuff that's really good at generating language,
00:39:09.712 Speaker 1 analyzing text, sort of formatting it the way you want it. It's a lot of the magic,
00:39:14.192 Speaker 1 but it's missing some of the, um, precision, maybe better reasoning, a sort
00:39:20.232 Speaker 1 of sequential, like if this, then that logic of a conversation. Um, it has a very
00:39:25.512 Speaker 1 short memory. Um, so that's some of the other limitations to be aware of. One final
00:39:30.352 Speaker 1 limitation to keep in mind is that this thing can only handle 4,000 tokens. So
00:39:36.332 Speaker 1 the GPT-3.5 model is limited in, like, its short-term working memory, how much
00:39:42.292 Speaker 1 it can know about. And the, the amount of space it has to speak back to you
00:39:48.232 Speaker 1 also uses that same shared amount of space. So 4,000 tokens is,
00:39:54.092 Speaker 1 is like 2,500. No, like 3,000 characters usually on average.
00:40:00.092 Speaker 1 So tokens are slightly different than, like, character size. You can, like, use a
00:40:05.772 Speaker 1 calculator to figure out, you know, how they get broken up. I think that's because
00:40:09.732 Speaker 1 some characters end up being condensed, um, down into one, but also white space and
00:40:15.512 Speaker 1 punctuation and stuff like that count. So you have roughly 3,000 characters of
00:40:20.692 Speaker 1 short-term working memory in your prompts. But that amount of space, and you could
00:40:24.812 Speaker 1 see here where I said that this was a learning opportunity, if you try to go above
00:40:28.872 Speaker 1 4,000, it will run out of space to answer you. So whatever you wanna tell this thing
00:40:33.252 Speaker 1 in terms of instructions, it has to fit within that 4,000. Um, so that's the only
00:40:38.552 Speaker 1 other gotcha to keep in mind. So from here, the way that this API works is
00:40:44.612 Speaker 1 you literally just send it this big blob of text, and it sends you back its answer.
00:40:49.252 Speaker 1 Now, in production, you're gonna use stuff like best of, and this is kind of like
00:40:55.992 Speaker 1 this part, right? Where you're trying to pick the best answer to send back. Generate
00:41:00.252 Speaker 1 five different answers and pick the best one. It's going to try to do something like
00:41:04.512 Speaker 1 that, okay? So, but it's more expensive to do that. So this is the general
00:41:10.291 Speaker 1 approach, and hopefully this was enough information that you guys can start
00:41:14.792 Speaker 1 prototyping your own ideas without a developer. Um, and you sort of feel out the
00:41:19.872 Speaker 1 boundaries of how hard is your problem. In this case, scam phone call detection
00:41:24.732 Speaker 1 looks very plausible. Uh, in fact, I would say that these state-of-the-art language
00:41:29.132 Speaker 1 models are exactly what we need to sort of curb this problem, and it's a very
00:41:33.972 Speaker 1 promising startup. And, you know, that is essentially my analysis for them is, yeah,
00:41:38.892 Speaker 1 this looks good. I think we can do this. And the most exciting part is you don't
00:41:42.472 Speaker 1 have to train your own models. This is just an off-the-shelf API from OpenAI. You
00:41:47.532 Speaker 1 can just prompt craft properly and then send it this information, and it acts as a
00:41:53.592 Speaker 1 scam phone call detection tool. Now, you're gonna need to do some extra steps to get
00:41:57.792 Speaker 1 to production to make the software ready. Obviously, there was the phone app and the
00:42:01.252 Speaker 1 other stuff, and there's gonna be some stuff you wanna do to make sure that it's not
00:42:04.832 Speaker 1 giving really bad answers. Maybe that's for another video. But when you're
00:42:08.572 Speaker 1 prototyping something, you don't care about any of that. All you really care about
00:42:12.472 Speaker 1 is, can I get this thing to work well enough that I can prove that it's possible?
00:42:18.472 Speaker 1 You can do that. The rest is basically normal software engineering. It's a solved
00:42:23.132 Speaker 1 problem. So there you have it. I hope that is helpful. Uh, I might prototype some
00:42:28.032 Speaker 1 more things, uh, for ChatGPT. It's not ChatGPT. It's just GPT-3.5. We will get the
00:42:33.572 Speaker 1 ChatGPT API hopefully in a couple weeks, maybe a couple months, and that's a whole
00:42:38.632 Speaker 1 new ballgame 'cause it's even better than this. But you will be far ahead of
00:42:42.732 Speaker 1 understanding how to prototype your ideas and what's possible if you get to work
00:42:46.912 Speaker 1 with just the black part, which, as you can see by the relative sizing, is the most
00:42:51.672 Speaker 1 important part. So don't delay. You can get started now and see what's possible.
00:42:58.832 Speaker 1 Hope that was helpful.
00:43:01.832 Speaker 1 So next time on Seeking Minima, we are going to explore a more
00:43:07.152 Speaker 1 philosophical fun question in the form of a story. So I wanted to try and take a
00:43:12.752 Speaker 1 stab at answering the question: what do we do when we don't have to do
00:43:18.852 Speaker 1 anything? And I tried to do that in a very entertaining way. Let me know what you
00:43:23.652 Speaker 1 think