All transcripts and search

Instructions For Prototyping ChatGPT Products

Uncorrected Scribe v2 machine transcript. Neutral speaker labels are specific to this recording. Review flags: 6.

TXT · SRT · VTT · Untouched JSON · Review flags

Source: youtube/20230212-XO-REklMsgc-Instructions For Prototyping ChatGPT Products.webm
SHA-256: f2bd0e6ede8a7c4ce8d5859525680db62c2013d0cc94a86943d7917c43a493d6

00:00:00.080 Speaker 1 [gentle music] This is Seeking Minimum.

00:00:10.460 Speaker 1 So today, we are going to take a look at how you can rapidly prototype some of your

00:00:14.740 Speaker 1 ChatGPT-related business ideas. Um, so what's really cool about this technique is

00:00:19.460 Speaker 1 that you don't have to be a software developer. You just have to be a little bit

00:00:23.220 Speaker 1 patient, and then you can sort of prove out whether ChatGPT is good enough or

00:00:28.820 Speaker 1 capable of solving your specific business problem. So, you know, let's just dive

00:00:33.600 Speaker 1 right into it. Uh, it's one of my favorite ways to rapidly prototype ideas, and

00:00:37.340 Speaker 1 today we're gonna walk through one of the actual ideas that I'm prototyping for a

00:00:40.840 Speaker 1 startup, and this is related around s-scam phone call detection, I guess is the

00:00:46.240 Speaker 1 right way to call it. Um, so but in your mind, you can walk through the steps for

00:00:51.040 Speaker 1 how you would prototype your ideas. So let's get started. So today, we are going to

00:00:55.540 Speaker 1 take a look at how you can rapidly prototype some of your ChatGPT-related business

00:01:00.300 Speaker 1 ideas. Um, so what's really cool about this technique is that you don't have to be a

00:01:04.540 Speaker 1 software developer. You just have to be a little bit patient, and then you can sort

00:01:08.860 Speaker 1 of prove out whether ChatGPT is good enough or capable of solving your specific

00:01:14.320 Speaker 1 business problem. So, you know, let's just dive right into it. Uh, it's one of my

00:01:18.640 Speaker 1 favorite ways to rapidly prototype ideas, and today we're gonna walk through one of

00:01:22.280 Speaker 1 the actual ideas that I'm prototyping for a startup, and this is related around

00:01:28.140 Speaker 1 s-scam phone call detection, I guess is the right way to call it. Um, so but in your

00:01:32.400 Speaker 1 mind, you can walk through the steps for how you would prototype your ideas. So

00:01:37.340 Speaker 1 let's get started.

00:01:48.560 Speaker 1 All right.

00:01:51.420 Speaker 1 Cool.

00:01:53.380 Speaker 1 So

00:01:55.140 Speaker 1 my name is Steven, and I like making things. I think that's enough of an

00:02:00.900 Speaker 1 introduction. So without further ado, here's what you're gonna do. So the first

00:02:06.360 Speaker 1 thing you're going to do is you need to go to beta.openai.com/playground.

00:02:12.760 Speaker 1 Um, this is separate from your, like, ChatGPT, like, uh, public user account. You're

00:02:17.560 Speaker 1 gonna wanna use a different account, and this is like a developer account. So what

00:02:23.140 Speaker 1 we're gonna do is we are going to use their DaVinci model. So

00:02:29.180 Speaker 1 you can see up here at the top it says Playground. Just click on that and click on

00:02:32.780 Speaker 1 text-davinci-003, 'cause that is the latest model.

00:02:37.420 Speaker 1 So like I said, I'm prototyping an idea for a startup that wants to be able to do

00:02:43.160 Speaker 1 scam phone call detection. So for the purposes of this exercise, what they want to

00:02:47.660 Speaker 1 do is they want to monitor, um, somebody's phone calls, outgoing and incoming, and

00:02:53.300 Speaker 1 basically continually listen to them and detect, is this a scam? Is this a scam?

00:02:59.220 Speaker 1 Is this a scam? And if it is, they're gonna make a decision. So that decision will

00:03:03.800 Speaker 1 end up being something like adding a trusted family member or friend to the phone

00:03:08.620 Speaker 1 call so that they can see what's going on, uh, or bringing in maybe a paid customer

00:03:12.760 Speaker 1 service agent who gets added to the call basically to verify, "Hey, who are you?

00:03:17.520 Speaker 1 What's going on?" You know, this person has a paid service to make sure that they

00:03:21.060 Speaker 1 don't get phone scammed. And the real use case here is that an enormous amount of

00:03:25.500 Speaker 1 money is scammed every year in the United States. Uh, s- like, something like $19

00:03:30.060 Speaker 1 billion, I think they said. Um, and most of those people are old people, and most of

00:03:35.640 Speaker 1 those people are the same subset of people over and over again. The nice way of

00:03:41.220 Speaker 1 putting this is that they are

00:03:44.140 Speaker 1 the trusting subset of our population. Um, so scam target-- scam victims

00:03:50.060 Speaker 1 are often, like, repeat victims. Um, so this is a really valuable product for a

00:03:55.780 Speaker 1 smaller portion of the population, but they need it really bad. So here's how this

00:04:00.740 Speaker 1 is gonna work, uh, and then we'll dive right into the prototyping stage, which

00:04:05.200 Speaker 1 anyone can do. So here, let's go here.

00:04:11.000 Speaker 1 Here's what we're gonna do. All right, so here's how you could imagine this app

00:04:13.580 Speaker 1 working, and maybe yours, you can do a similar step in your own thinking,

00:04:19.779 Speaker 1 right? So don't think just because you're not a software developer, you can't think

00:04:23.180 Speaker 1 this through. Like, it's not that hard. So imagine

00:04:29.040 Speaker 1 this, right? You got a smartphone, right? Somebody pays for the service. However,

00:04:32.680 Speaker 1 it's X dollars per month, and what happens is they have an app running on their

00:04:38.300 Speaker 1 phone. So

00:04:42.040 Speaker 1 let me move this down a little bit. They're gonna have an app running on their

00:04:45.300 Speaker 1 phone, so they have an app, right? This app is going to listen to their calls,

00:04:51.900 Speaker 1 and it's going to send

00:04:54.640 Speaker 1 the audio data

00:05:00.080 Speaker 1 to our servers. All right? It's gonna send them to our servers. Do,

00:05:05.980 Speaker 1 do-do, do-do.

00:05:09.780 Speaker 1 But in this case, we didn't even have to do that. We're gonna send this to

00:05:15.700 Speaker 1 Whisper.

00:05:18.900 Speaker 1 Whoop.

00:05:20.980 Speaker 1 And this is gonna be... So Whisper is another OpenAI API which turns

00:05:27.240 Speaker 1 audio into text. It's transcription software. It's very good. Um, when I said we

00:05:32.700 Speaker 1 could do this on our own servers, Whisper is actually open source. You don't have to

00:05:36.640 Speaker 1 pay for it, but you can. So bit of confusion there. So then

00:05:42.340 Speaker 1 Whisper is going to turn the audio into text,

00:05:47.440 Speaker 1 right? So you can just sort of walk through a similar pipeline in your mind, and

00:05:52.540 Speaker 1 then the text is going to be sent to...

00:05:57.680 Speaker 1 We'll just label this GPT for now. It's going to be sent to a language model, and

00:06:02.620 Speaker 1 the language model is gonna do some thinking.

00:06:10.250 Speaker 1 Thinking, right? And then from there, it's gonna go to our servers.

00:06:16.190 Speaker 1 It's going to decide scam,

00:06:20.050 Speaker 1 not scam, right? And then from there, we'll make some more decisions, right? So

00:06:26.290 Speaker 1 if what you wanna test is: Can my product idea work? Is this part,

00:06:32.670 Speaker 1 the AI part, is it good enough to do what I want to do? That is what today's video

00:06:37.770 Speaker 1 is about, and that is the part we're going to test. The rest of this, we're all

00:06:41.270 Speaker 1 going to assume is basic software development, and it can be done. If the hard part

00:06:45.550 Speaker 1 works, that's very promising for the rest of it. All right. First, we're gonna get

00:06:49.290 Speaker 1 into building out the prototype a little bit, and then we'll come back to the

00:06:52.150 Speaker 1 whiteboard to show you a little bit about why this works and how you can think about

00:06:57.030 Speaker 1 the capabilities of these language models.

00:07:00.230 Speaker 1 Okay. So what we want here for the blue part is, okay, basically a

00:07:06.070 Speaker 1 transcription of a call log, essentially. You know, call starts, back and forth,

00:07:10.610 Speaker 1 kinda like a chat log. Person A says something, person B responds. Person A says

00:07:14.669 Speaker 1 something, person B responds. So we want this thing to continually analyze the

00:07:19.930 Speaker 1 transcript from the very start all the way to the end.

00:07:25.970 Speaker 1 Hi, this is Stan. Right. And we want this language model to evaluate

00:07:32.030 Speaker 1 line by line in the transcript, is this a scam? Is this a scam? Is this a scam?

00:07:38.290 Speaker 1 So it's taking all the context of what has happened before this into account, but

00:07:43.250 Speaker 1 it's also doing it in real time, right? Because it's not useful to go 20 minutes

00:07:47.930 Speaker 1 down a conversation, send the whole recording to somebody and say, "Oh, yeah. Yep,

00:07:53.430 Speaker 1 you definitely got scammed. [chuckles] That was a scam. Sorry." Uh, but... So you

00:07:58.970 Speaker 1 wanna do it in real time, right? Because you wanna step in before it's too late. So

00:08:03.910 Speaker 1 that's what we're going to try to build. Um, so it needs to know as soon as possible

00:08:07.950 Speaker 1 whether it's confident it's a scam or not. So how are we gonna do that?

00:08:12.610 Speaker 1 Well, this might be kinda surprising if you haven't spent much time with ChatGPT,

00:08:16.810 Speaker 1 um, but the way you do that [chuckles] is like this. You

00:08:22.670 Speaker 1 are a scam detection tool.

00:08:29.730 Speaker 1 You will review a call, live call

00:08:36.150 Speaker 1 transcript and decide whether

00:08:42.050 Speaker 1 our user is talking to a scammer or not.

00:08:49.410 Speaker 1 I don't know about you, but this is profound that this is how this works. It's

00:08:52.690 Speaker 1 insane. This is literally what you do, okay? So it takes the same format

00:08:58.690 Speaker 1 every time, right? When you're building these things out. Roughly speaking, it's in

00:09:02.370 Speaker 1 three parts. Forgive me, I'm gonna use this section as notes for now. All right. You

00:09:06.570 Speaker 1 have part one,

00:09:09.330 Speaker 1 high level direction and role. So the language model needs

00:09:15.230 Speaker 1 to know what role it's playing. Then you have part two,

00:09:22.190 Speaker 1 in

00:09:24.790 Speaker 1 bounds example.

00:09:28.570 Speaker 1 Okay. So this is giving it an actual transcript. So w- first,

00:09:34.830 Speaker 1 you tell it what this is and what role it's playing. Then you're going to tell it,

00:09:39.470 Speaker 1 "Here is an example of a scam call, okay? And how I want you to

00:09:45.050 Speaker 1 respond." Right? You show it. You... First, you tell it what your role is, so it can

00:09:48.750 Speaker 1 constrain itself, so it knows roughly how to respond, what it's supposed to be

00:09:53.310 Speaker 1 doing. And then you give it an example, so it knows exactly how to format exactly

00:09:57.630 Speaker 1 what you want its responses to be. [lip smack] If you've played with ChatGPT, you'll

00:10:01.210 Speaker 1 already understand this so far. Then you want to give it

00:10:06.590 Speaker 1 an out of bounds example.

00:10:12.750 Speaker 1 All right. So I guess one way of framing this is that you want a true positive,

00:10:20.290 Speaker 1 okay? And then you want a true negative. So true negative, this is a scam call. Then

00:10:25.810 Speaker 1 you need to say, "Okay, here's an example of a conversation that is not a scam call,

00:10:30.650 Speaker 1 so that you can tell the difference. Here's a pushy collections department for

00:10:35.090 Speaker 1 medical billing," right? That's not a scam call. But they actually have a lot of the

00:10:38.770 Speaker 1 same qualities, so you need to provide some way for it to understand the distinction

00:10:42.890 Speaker 1 between them. And then, I don't know, maybe you include an example of something like

00:10:47.610 Speaker 1 a charity. You know, they're calling and asking for your help, but they do want your

00:10:50.830 Speaker 1 money, but it's not really a scam. How do you do it-- How can you tell the

00:10:54.790 Speaker 1 difference? Or somebody trying to upsell you on something.

00:10:59.250 Speaker 1 All right. So part four.

00:11:05.150 Speaker 1 Please don't mind my children screaming in the background. Can't turn 'em off.

00:11:09.490 Speaker 1 Finally, you wanna provide an out of domain example. This is like

00:11:14.310 Speaker 1 [chuckles] a call that is, like, totally left field. Like, it, it is not just a scam

00:11:20.510 Speaker 1 or, like, um... So this is more like false

00:11:26.270 Speaker 1 positive, right? Things you think that it will accidentally classify as a scam.

00:11:32.410 Speaker 1 And then you need to give out of domain example. So this is, like, true negative.

00:11:37.770 Speaker 1 This is, like, somebody talking to somebody else about, I don't know, their,

00:11:42.730 Speaker 1 [chuckles] their favorite baseball players. Two friends who know each other well,

00:11:45.810 Speaker 1 right? It's absolutely nothing to do with money or scamming at all, so that this

00:11:50.630 Speaker 1 thing can know where the boundary is for when it should be paying attention or not,

00:11:55.350 Speaker 1 right? So it should be paying attention in this example, because it is about money.

00:11:59.890 Speaker 1 It is about people trying to extract value from you and trying to coerce you, but

00:12:04.330 Speaker 1 it's not a scam. Um, so it needs to be paying attention here. And in this example,

00:12:10.210 Speaker 1 you know, just tune out. These people are talking about baseball. Like, you don't

00:12:13.210 Speaker 1 need to worry about this. All right. So that's roughly what you're gonna do, and

00:12:16.570 Speaker 1 you're gonna provide as many of these as you can Within reason. We'll talk about it.

00:12:21.634 Speaker 1 So let's fill this in. So this is part one. Part two will be something like, um...

00:12:27.194 Speaker 1 So you also need to give it--

00:12:30.894 Speaker 1 That's part one. So now we need to give it some high-level direction on the

00:12:35.134 Speaker 1 formatting. So you will add annotations to each line of the

00:12:40.874 Speaker 1 transcript

00:12:43.134 Speaker 1 and indicate whether this line contains clues

00:12:49.314 Speaker 1 that increase your suspicion

00:12:54.974 Speaker 1 of it being a scammer or a scam call, or decrease your suspicion.

00:13:01.594 Speaker 1 This is one of the beautiful things about language models, is they can explain

00:13:04.174 Speaker 1 themselves. So we want it not just to, hey, at some point, just randomly interrupt

00:13:09.774 Speaker 1 and say it's a scam phone call. You would like it to be able to say line by line,

00:13:13.574 Speaker 1 "Ooh, that was suspicious. Eh, that's not suspicious. Actually, that's suspicious."

00:13:18.974 Speaker 1 So that it has, like, this increasing or decreasing, like, thermometer of how scammy

00:13:23.514 Speaker 1 this is or isn't, and then at some point, it should cert- short circuit and say,

00:13:27.234 Speaker 1 "Okay, this is definitely a scam phone call." All right. So that's what we're gonna

00:13:30.694 Speaker 1 emulate. So now we want it to give it line by line. Here's an example

00:13:38.014 Speaker 1 transcript with your annotations

00:13:43.994 Speaker 1 after the pipe character. There's just an arbitrary way of formatting it. So let's

00:13:49.374 Speaker 1 go grab an example conversation. All right, so I've already typed this up. "Hey,

00:13:55.514 Speaker 1 here's what I want you to do." So let's delete that for now.

00:14:00.094 Speaker 1 See this? Do it like that. How insane is that? "Hey, you, computer,

00:14:05.834 Speaker 1 [chuckles] do it like this." [laughs] It just blows my mind. Um,

00:14:11.974 Speaker 1 understanding, uh, a fair amount of the technicals behind it does not make it less

00:14:15.574 Speaker 1 mysterious. It really is insane that you can do this. So this is not retraining the

00:14:21.314 Speaker 1 neural networks, uh, under these language models. Um, that's a fancy way of saying

00:14:26.594 Speaker 1 it already knows how to do stuff like this, and you're just giving it basic

00:14:29.894 Speaker 1 instructions. Uh, rather than thinking of this like

00:14:35.734 Speaker 1 having to learn a new skill, it doesn't have to learn a new skill. It already knows

00:14:39.114 Speaker 1 how to do this. You're just having to give it instructions on what you want, like

00:14:42.954 Speaker 1 teaching a person how to play a new game. They already know how to pick up a ball

00:14:46.194 Speaker 1 and throw it in a hoop, but if there are new rules about what makes a scoring point,

00:14:50.834 Speaker 1 what is a penalty, they have to learn those things. But that's not really a new

00:14:54.694 Speaker 1 skill so much as they just need to understand the parameters. So it's crazy. So

00:15:00.274 Speaker 1 here's what we want it to do, right? Here's an example call. So this-- The call

00:15:04.094 Speaker 1 starts. Our user is Harold. "Hi, is this Harold?" So I use the little up caret for

00:15:10.054 Speaker 1 increased suspicion.

00:15:12.754 Speaker 1 Is

00:15:14.614 Speaker 1 increased suspicion. And then I'll use a V for, like, a down caret, is decreased

00:15:21.274 Speaker 1 suspicion.

00:15:27.054 Speaker 1 Then

00:15:30.974 Speaker 1 a line is not important.

00:15:34.854 Speaker 1 You may enter white space,

00:15:39.314 Speaker 1 single space.

00:15:42.114 Speaker 1 Uh, and then otherwise, add a very brief

00:15:48.514 Speaker 1 note explaining

00:15:54.554 Speaker 1 why this clue is relevant, right?

00:15:58.594 Speaker 1 So here we go. So basically, you add in your prefilled example. So here we go.

00:16:03.814 Speaker 1 Somebody calls Harold. "Hi, is this Harold?" I decided, you know what? If the caller

00:16:08.954 Speaker 1 doesn't already know the person they're calling,

00:16:12.374 Speaker 1 that's a minor suspicion point. All right? So let's go from here. "Yeah, this is...

00:16:18.074 Speaker 1 Yes, speaking." Anyway, somebody from Doggone [chuckles] is calling. "How you doing

00:16:23.354 Speaker 1 today?" Uh, they don't have any rapport with this user at all. They have to

00:16:26.294 Speaker 1 establish some rapport. Hmm. Scammers, by and large, are strangers. Uh, so let's

00:16:32.054 Speaker 1 keep going. And this is where having knowledge of your domain is actually very

00:16:35.634 Speaker 1 useful. There's a lot of things that I can input here that indicate whether this is

00:16:40.434 Speaker 1 a scam call or not. So, you know, off the cuff, that will be, uh, trying to get the

00:16:45.674 Speaker 1 buyer to, uh, um, the caller to convert on that actual call, as in they can't

00:16:51.654 Speaker 1 let them go. They have to transfer some sort of value right here on the call or

00:16:55.574 Speaker 1 provide information. Um, some other clues are the scammer is asking--

00:17:01.654 Speaker 1 has a little bit of information about them, but is asking for more pertinent or,

00:17:06.114 Speaker 1 like, more specific information, like a Social Security number or a home address.

00:17:11.194 Speaker 1 Um, so they, they have a breadcrumb, and they're asking you for the loaf, right? So,

00:17:16.474 Speaker 1 so they can use that, uh, for identity theft or other stuff. Um, they're trying to

00:17:20.774 Speaker 1 get, um, the user to buy gift cards. Uh, gift cards are a way of essentially

00:17:26.754 Speaker 1 laundering money, [chuckles] right? Um, or prepaid credit, credit cards. Uh, that's

00:17:31.854 Speaker 1 another common tactic. There's lots of, like, through lines in these scams that you

00:17:36.374 Speaker 1 can use those as red flags for very strong signals, and we would use those in our

00:17:40.354 Speaker 1 example calls. So anyway, if you read through this, um, what you can see is that

00:17:45.314 Speaker 1 this is basically a pushy,

00:17:48.654 Speaker 1 uh, a pushy charity caller. Um, and then at the end, I short circuit, and I want

00:17:54.594 Speaker 1 it to output this part below the lines. So analysis, not a scam, 18% chance, right?

00:18:00.794 Speaker 1 So, and tell me why. Again, crazy thing about language models. Oh, yeah. If you

00:18:04.834 Speaker 1 decide this is not a scam at any point, I want you to just tell me why. Uh, and here

00:18:09.574 Speaker 1 it does. So it says, you know, they were pushy. They were trying to collect money,

00:18:14.354 Speaker 1 which is a warning sign, but it's not conclusive because it seems like Harold has

00:18:17.814 Speaker 1 already donated to them be-before because he did not deny dono-donating when this

00:18:22.394 Speaker 1 person said that he did, right? So- That's the crazy thing about these things, is

00:18:27.646 Speaker 1 they can do actual common sense logic. So what I'm gonna do now

00:18:33.646 Speaker 1 is I'm actually going to input some prompts that I had from before, and we're gonna

00:18:38.426 Speaker 1 just dump these in, 'cause I'm not gonna retype all this stuff. So here you can see,

00:18:43.966 Speaker 1 uh, my earlier prompt. I did a very short example here. And let's just see what this

00:18:49.566 Speaker 1 thing does, and I'm gonna show you how to work with it. So I've got a bunch of part

00:18:53.706 Speaker 1 one, two, three, four, you know, some examples in here so that this thing

00:18:57.666 Speaker 1 understands its role in the conversation. And now we're gonna come down here, and

00:19:02.486 Speaker 1 we're going to try it out. So let's start a new call. Um, so

00:19:08.406 Speaker 1 let's make sure we have the same format every time. So here we go. Let's start a new

00:19:12.706 Speaker 1 live phone call.

00:19:15.586 Speaker 1 Just copy this.

00:19:19.386 Speaker 1 Imagine that this is where the real entry hits the, the OpenAI servers, right?

00:19:26.086 Speaker 1 To be analyzed by, in this case, their DaVinci model. We'll get into what the

00:19:30.106 Speaker 1 difference is between this DaVinci model and ChatGPT in a moment, and why that

00:19:34.306 Speaker 1 matters for when you're prototyping stuff.

00:19:38.266 Speaker 1 Okay. So

00:19:43.186 Speaker 1 outbound call. So let's make this an inbound call.

00:19:47.266 Speaker 1 Inbound call.

00:19:50.166 Speaker 1 What is it? It says 239.

00:19:53.706 Speaker 1 Uh, user is Tony.

00:19:56.686 Speaker 1 Okay.

00:19:58.806 Speaker 1 So the caller says,

00:20:01.126 Speaker 1 "Hi,

00:20:06.866 Speaker 1 is Mr.,

00:20:12.306 Speaker 1 uh,

00:20:14.346 Speaker 1 Tony Terrence..."

00:20:16.766 Speaker 1 The real name.

00:20:17.506 Speaker 1 [chair creaking]

00:20:31.326 Speaker 1 There we go. "Uh, is Mr. Tony Terrence available?"

00:20:35.986 Speaker 1 All right, so now we put a pipe, and this is the part where

00:20:40.766 Speaker 1 we come in here, and we're going to need to enter our stop sequences and the start

00:20:45.446 Speaker 1 text. So when the m- language model sees this character,

00:20:51.746 Speaker 1 we want it to stop

00:20:55.346 Speaker 1 talking. All right. So, well, let me show you what happens if you don't do this.

00:20:58.846 Speaker 1 Right? Here. Hand all of this t- off to the language model. So we've given it its

00:21:03.166 Speaker 1 instructions. We've given it lots of examples, and now we've given it the beginning

00:21:06.506 Speaker 1 of a real call.

00:21:11.406 Speaker 1 So you can see what it does. [chuckles] It tries to actually fill in

00:21:17.426 Speaker 1 the entire call, uh, which we don't want it to do. So [chuckles] see, look, it,

00:21:23.526 Speaker 1 it's just trying to make up a scam phone call conversation, which is hilarious, uh,

00:21:27.906 Speaker 1 but not what we're looking for. What we really want it to do is, "Hey, when we enter

00:21:31.546 Speaker 1 a new line, meaning this person has finished speaking, stop talking. Because what I

00:21:36.566 Speaker 1 want you to generate is I want you to generate the analysis. Don't try to make up a

00:21:41.946 Speaker 1 conversation." So press the enter button for the new line. That means stop talking

00:21:47.266 Speaker 1 after you hit a new line. Um, and that should be sufficient. So let's clear this

00:21:52.046 Speaker 1 out.

00:21:54.106 Speaker 1 All right, so now let's try again. "Hi, is Mr. Tony Terrence available?" Imagine

00:21:58.266 Speaker 1 this is streaming live. Right?

00:22:01.786 Speaker 1 Cool. See how it stops? "Caller does not know user." It already knows what to do.

00:22:07.606 Speaker 1 [chuckles] This is the craziest thing in the world to me. Okay, so imagine this is a

00:22:11.386 Speaker 1 live call happening, and our user, Tony, is receiving this call, right? And it's our

00:22:15.946 Speaker 1 job for our bot to properly analyze this scun- scam phone call or not, and also

00:22:21.886 Speaker 1 not to annoy Tony if this is not a scammer. All right, so here we go. Tony says... I

00:22:27.926 Speaker 1 don't know what he's gonna say.

00:22:30.806 Speaker 1 Says, "Tony, who's this?"

00:22:35.066 Speaker 1 Okay.

00:22:41.486 Speaker 1 Cool. Space. That's what we want it to do. That line is not important. Um, you know,

00:22:47.286 Speaker 1 maybe he would've said that, or maybe it would've said that, "Caller doesn't know

00:22:50.026 Speaker 1 user," but it's already said that, so there's no reason to do that. Um, by the way,

00:22:53.626 Speaker 1 this dial over here, the presence and frequency penalties, if you find that it's

00:22:57.426 Speaker 1 offering the same analysis over and over again, and it's not actually useful, you

00:23:00.866 Speaker 1 can increase these sliders, and it will stop repeating itself so much, and it will

00:23:06.446 Speaker 1 try to find new answers. So that's worth doing.

00:23:10.586 Speaker 1 All right, so I don't know. Let's just finish up this transcript real quick so that

00:23:13.566 Speaker 1 we can get back into the theory of, uh, you know, how to think about prototyping

00:23:18.286 Speaker 1 your products. All right, so this is Tony. "Who's this?" Um,

00:23:23.986 Speaker 1 uh, "This is Megan with the IRS.

00:23:29.766 Speaker 1 I have some serious matters to discuss with you.

00:23:44.646 Speaker 1 I'm a tax collections agent

00:23:50.846 Speaker 1 with the IRS," right?

00:23:54.086 Speaker 1 It will alw- almost always appeal to some kind of authority.

00:24:01.546 Speaker 1 Oh, so here we go. Let's see what it does with that.

00:24:10.906 Speaker 1 Hold on. "Caller

00:24:16.926 Speaker 1 is trying to scare user." Interesting analysis. Okay. I suppose that that

00:24:22.946 Speaker 1 is an intimidating opening line. [chuckles] All right, let's see where it goes from

00:24:26.506 Speaker 1 here. Uh, Tony says

00:24:39.216 Speaker 1 We already paid those.

00:24:45.676 Speaker 1 Thought.

00:24:47.616 Speaker 1 Let's, uh,

00:24:49.776 Speaker 1 do that. See what it does with that.

00:24:53.156 Speaker 1 Probably nothing.

00:24:55.936 Speaker 1 User's trying to resolve an ex-existing issue. See, this is not useful. We

00:25:01.816 Speaker 1 don't want it to do that. It doesn't make any sense. They're trying to resolve an

00:25:06.196 Speaker 1 existing issue. I mean, I guess you could see it that way because somebody

00:25:11.936 Speaker 1 called sort of trying to intimidate them, and our user, Tony, has responded

00:25:17.876 Speaker 1 by saying that there is an existing issue that is related. So because they have a

00:25:22.376 Speaker 1 past case with the IRS, this software thinks that, oh, well, just like we

00:25:28.416 Speaker 1 showed it in that initial example,

00:25:31.876 Speaker 1 the charity, because this person has associated with this charity in the past, they

00:25:35.696 Speaker 1 have established trust with them. They probably do not... This is probably not a

00:25:41.055 Speaker 1 scam. Okay? So let's keep going and see where it goes.

00:25:45.676 Speaker 1 This is an opportunity maybe to correct part of the bounds of where it will get

00:25:49.536 Speaker 1 things wrong.

00:25:51.856 Speaker 1 Um, okay. Caller. So this is our Megan, the IRS agent. Uh.

00:26:04.896 Speaker 1 They're gonna cite existing exact numbers usually, um, because exact numbers sound

00:26:10.076 Speaker 1 like you know what you're talking about and you have a real document in front of

00:26:12.316 Speaker 1 you.

00:26:15.116 Speaker 1 Three hundred and forty and 22 cents.

00:26:24.096 Speaker 1 All

00:26:38.736 Speaker 1 right. So they're gonna continue the intimidation techniques. This is often what

00:26:42.716 Speaker 1 happens.

00:26:45.336 Speaker 1 So it thinks that line [chuckles] is not relevant. I don't know about that. We'll

00:26:49.196 Speaker 1 get back to that.

00:26:55.096 Speaker 1 So let's see what it does here.

00:26:59.536 Speaker 1 And then at some point, I'm gonna short circuit and tell it, "Give me your

00:27:01.676 Speaker 1 analysis."

00:27:06.676 Speaker 1 So now it's thinking, because this guy didn't get any letters, maybe this isn't an

00:27:11.396 Speaker 1 existing issue.

00:27:53.856 Speaker 1 So they're gonna try to elicit sympathy while also sort of issuing a veiled threat.

00:27:57.236 Speaker 1 "You owe this money right now."

00:28:04.096 Speaker 1 Caller is in a rush, is trying to rush user. Remember, we told it early-- Maybe you

00:28:09.316 Speaker 1 don't, but early on, I told it that if the caller is in a... Basically, they're

00:28:14.636 Speaker 1 trying to convert quickly, then they're probably likely to scam them. Okay.

00:28:20.676 Speaker 1 So anyway, so this is just a fake conversation. So let's go ahead and end the call

00:28:26.136 Speaker 1 and give me your analysis.

00:28:34.316 Speaker 1 Possible scam, 75% chance. Okay. Final reasoning. Tell me why you

00:28:40.276 Speaker 1 think it's a scam, right?

00:28:47.736 Speaker 1 Mm. Hold on. There's a separate problem, which is, is instructional but not

00:28:53.416 Speaker 1 exciting. So

00:28:56.555 Speaker 1 all right, let's do that.

00:29:02.396 Speaker 1 I think it's 'cause there's a space there.

00:29:08.456 Speaker 1 Stop sequence. How is the... Oh, it was formatted like that. Mm. Let me remove the

00:29:13.076 Speaker 1 stop sequence. It's because I was putting this on a new line. It's hitting a new

00:29:17.696 Speaker 1 line and then giving up. Um, if I had formatted all the prior examples like that,

00:29:22.316 Speaker 1 [chuckles] this would've worked. Just a silly little quirk. Give me a second. Uh,

00:29:27.276 Speaker 1 try again.

00:29:30.456 Speaker 1 See, there it goes, 'cause it's trying to put it on the next line. All right. So

00:29:33.136 Speaker 1 what does it have to say about this phone call? Caller's claiming to be a tax

00:29:36.096 Speaker 1 collections agent with the IRS and trying to scare the user with serious matters.

00:29:39.336 Speaker 1 Tony is trying to resolve an existing issue and hasn't received any cor- official

00:29:42.456 Speaker 1 correspondence. Caller's also trying to rush them, which is common trait of scam

00:29:46.696 Speaker 1 phone calls. Eh, it's decent analysis. So a lot of getting these things right is

00:29:51.936 Speaker 1 going back and seeing where they break. They have common sense. When you think about

00:29:56.676 Speaker 1 these mo- large language models, especially like GPT,

00:30:00.896 Speaker 1 they have common sense. Think of them as like commodified intelligence. They have

00:30:04.316 Speaker 1 the un- the basic intelligence of like a 20-year-old, right? They have some common

00:30:09.976 Speaker 1 sense, um, and they know a lot of things, but they're not really

00:30:15.716 Speaker 1 necessarily expert. Um, so you can think of them as like a common sense reasoning

00:30:20.376 Speaker 1 machine, but you do have to show them the patterns. Um, s- and there are places

00:30:24.976 Speaker 1 where their reasoning will break down, especially if you've given them prior

00:30:27.996 Speaker 1 examples. They will overweight those examples sometimes. So in this case, I think a

00:30:33.656 Speaker 1 better way to fix this, and this is how I'm improving this prototype, is to You go

00:30:39.316 Speaker 1 back over, you do these little hand-weighted tests, like, "Okay, I did a fake co-

00:30:43.176 Speaker 1 phone call to see what your reasoning is at every line." So you're trying to scare

00:30:46.815 Speaker 1 them, that makes sense. Trying to... So my, my recording cut out there. All right.

00:30:52.476 Speaker 1 So anyway, yeah, they're, they're trying to scare them, that made sense. But what I

00:30:56.476 Speaker 1 did is I went in and I actually edited the, the responses from the AI as if I were

00:31:02.356 Speaker 1 the AI, so sort of just hand-correcting it. And then this is how it can really learn

00:31:06.896 Speaker 1 the nuances of the boundaries of your problem. And you just go through in this

00:31:10.316 Speaker 1 iterative cycle, and you, you see what it says. If it does, says something that

00:31:14.136 Speaker 1 doesn't make sense, you go back and correct it, and then you go again, and then

00:31:17.316 Speaker 1 eventually it's pretty good. Sometimes you will have to come back in here and change

00:31:21.916 Speaker 1 your original prompt, and that will make a big difference, right? So here I might

00:31:27.576 Speaker 1 add something about, you know, gift cards being a hallmark of scam phone callers or

00:31:31.876 Speaker 1 some of the other nuances of a conversation.

00:31:35.596 Speaker 1 Okay. Some other things to keep in mind.

00:31:38.896 Speaker 1 This is not ChatGPT. What it is, is sort of like what's under ChatGPT's

00:31:44.916 Speaker 1 hood. So it is a large language model called GPT-3.5.

00:31:50.976 Speaker 1 That's the version. Specifically, it's DaVinci 03. Okay. So you can see all the

00:31:56.336 Speaker 1 different versions of the GPT language model here. There's one for code, so

00:32:01.676 Speaker 1 technically this DaVinci to Code DaVinci is also under ChatGPT. Th-this is the part

00:32:06.516 Speaker 1 of ChatGPT that knows how to write code. This is the part that knows how to write

00:32:10.096 Speaker 1 language and other stuff. So what this actually is, is not the full

00:32:16.336 Speaker 1 ChatGPT, right? This is just part of ChatGPT.

00:32:22.216 Speaker 1 So let's go ahead and start over here.

00:32:28.876 Speaker 1 So, and this is worth explaining because the ChatGPT API is not out yet. You know,

00:32:34.776 Speaker 1 it will be out hopefully soon, couple weeks, maybe a couple months. Um, but you can

00:32:38.856 Speaker 1 prototype before that, and you can do it without having ChatGPT, because you still

00:32:44.576 Speaker 1 have a big piece. So what this is, ChapGT- ChatGPT, as far as I underthis-

00:32:48.776 Speaker 1 understand the architecture, is in three pieces, right? We have the part we just

00:32:53.476 Speaker 1 talked about, so this is your DaVinci.

00:32:57.916 Speaker 1 And there are other ones, as you saw. This is the GPT-3.5 language model. This is

00:33:03.476 Speaker 1 kinda like the engine in your car. This is the part that is just like, has all

00:33:09.436 Speaker 1 the information, knows how to generate language. For the most part, this is the part

00:33:13.976 Speaker 1 that knows how to generate any speech. Really, it's a big fancy autocomplete.

00:33:20.176 Speaker 1 If I start typing that, it will know

00:33:26.736 Speaker 1 how to finish it.

00:33:29.276 Speaker 1 Thing here, then I... Right? So what this is, is a big fancy autocomplete.

00:33:35.676 Speaker 1 That's it. That's kind of what this is. But it's, it's a very good one. Um,

00:33:41.896 Speaker 1 and it's most of the magic behind ChatGPT. Now, ChatGPT, so this is what we're

00:33:46.536 Speaker 1 using, right? Is the big fancy autocomplete. It has a lot of the common sense

00:33:50.516 Speaker 1 reasoning abilities that ChatGPT has. It's very good, but it's not quite as good,

00:33:56.636 Speaker 1 and I wanted you to understand why. So ChatGPT

00:34:02.536 Speaker 1 has two other pieces sitting on top of it. It has a

00:34:08.836 Speaker 1 reinforcement [chuckles] There we go.

00:34:15.296 Speaker 1 Model, and it also has... Here we go.

00:34:20.536 Speaker 1 Forgive the words bleeding over into one another. It also has a reward model.

00:34:28.656 Speaker 1 Model. And just in layman's speak, what this is, a reinforcement learning model...

00:34:34.476 Speaker 1 Oh, sorry, reinforcement learning. Forgot that word. What this is, is it's kind of

00:34:39.416 Speaker 1 like a, uh, video game AI agent. It has a state, its understanding of

00:34:45.416 Speaker 1 the world. So in this case, its state is the conversation so far. It has an

00:34:49.716 Speaker 1 understanding of... And included in that are, like, other players in the world. It

00:34:54.795 Speaker 1 has a set of actions that it can take, and then it has basically a way of predicting

00:35:01.136 Speaker 1 or deciding how to make its move. So you can think of it as like, a reinforcement

00:35:06.676 Speaker 1 learning model is a kind of AI that takes a look at the world,

00:35:11.496 Speaker 1 takes a look at what actions it is allowed to take on this one time step, and then

00:35:16.156 Speaker 1 picks one action to take, and then it sees how the world updates.

00:35:22.156 Speaker 1 Okay, so you can imagine this pretty easily for a video game. Like if it's

00:35:25.436 Speaker 1 tic-tac-toe, it's watching the state of the board. That's its state of the world. It

00:35:29.276 Speaker 1 knows it can place one of its letters in one of the squares, so it will decide

00:35:35.296 Speaker 1 where to put it, then it will take an action, put the letter there, and then it will

00:35:38.976 Speaker 1 wait to see what the other player does, and then it will take another look at the

00:35:42.316 Speaker 1 state of the world, decide what to do next. So it just runs in a loop like this. See

00:35:46.736 Speaker 1 what's happening, see what I can do, choose one of the things I can do, and sort of

00:35:51.896 Speaker 1 predict what effect it will have, and to get my reward, right? So it's trying to

00:35:56.496 Speaker 1 take, get closer and closer to its reward at every step. So there's this thing,

00:36:02.816 Speaker 1 so that's cool. Uh, that's one of the things where this help basically gives it

00:36:08.476 Speaker 1 greater context. It makes it a better chatbot, um, since it's trying to

00:36:14.156 Speaker 1 predict the answer that you want, right? So you can think of it like the state of

00:36:20.056 Speaker 1 the world is the conversation, in this case. This is for ChatGPT. And it's trying to

00:36:24.956 Speaker 1 predict the right answer, and it has this big old, uh, text engine

00:36:30.936 Speaker 1 underneath the hood that it's basically poking and prodding, just like we were over

00:36:35.416 Speaker 1 here with the prompts. It's doing something like that and saying, "Hey, [chuckles]

00:36:39.876 Speaker 1 give me a bunch of answers out, and I'm going to try to pick the answers that look

00:36:44.396 Speaker 1 like what this person over here wants." Right? So it's poking and prodding at

00:36:50.352 Speaker 1 the engine under the hood just like we are. However, it has some other tools to

00:36:55.252 Speaker 1 figure out what's a good answer, and this is another one of ChatGPT's innovations.

00:36:59.932 Speaker 1 The reward model is human-taught. So I think it's a

00:37:05.572 Speaker 1 human reinforced learning. I can't remember. Human behavior modeling, something like

00:37:10.332 Speaker 1 that. So I think the gist is when they're training ChatGPT, what they do is

00:37:16.632 Speaker 1 they have a bunch of human-asked questions and human responses. Okay?

00:37:22.732 Speaker 1 And then slowly, they start mixing in these two pieces, you know, as they come up

00:37:28.372 Speaker 1 with their own answers. So this thing is learning on that, and throughout the

00:37:32.472 Speaker 1 training, they're going to start mixing in these answers with the human answers.

00:37:38.632 Speaker 1 Right?

00:37:43.092 Speaker 1 And a human, a real live human, is going to pick which one is best. Okay,

00:37:49.012 Speaker 1 so this is like, you know, like Amazon Machine Turk or whatever it's called. Like,

00:37:53.572 Speaker 1 this is literally just like people coming in [chuckles] and just like, "All right,

00:37:57.592 Speaker 1 this person asked this question. Which of these five answers is the best answer?"

00:38:01.572 Speaker 1 And real humans will pick. It's kinda like the Turing test. And eventually, they're

00:38:05.512 Speaker 1 gonna start mixing in the AI answers, and what you hope is that the AI answers

00:38:10.332 Speaker 1 either are as good or picked as often as the human answers, or they're picked more

00:38:15.252 Speaker 1 often than the human answers. So eventually, what you've trained is a little reward

00:38:19.512 Speaker 1 model that is pretty good at picking the kinds of answers that humans pick. That's

00:38:24.652 Speaker 1 what it does. It's literally just like a little brain, and all this brain does...

00:38:28.992 Speaker 1 [chuckles] That was my drawing for brain. All this brain does is it knows how to

00:38:34.132 Speaker 1 pick answers that look like the answers that a human would pick. So using all of

00:38:38.632 Speaker 1 these pieces together, so, hey, it's going to use this thing to evaluate whether the

00:38:44.552 Speaker 1 output of this thing is any good, and it's gonna generate a bunch of them. Oh,

00:38:49.252 Speaker 1 sorry, I should've used black since that was the AI color. It's gonna generate a

00:38:52.712 Speaker 1 bunch of them,

00:38:55.412 Speaker 1 send them into this thing, have it pick the best one, and then give that one back to

00:39:00.032 Speaker 1 you. So it's missing these two pieces. So what it is is the

00:39:05.972 Speaker 1 big fancy autocomplete, the stuff that's really good at generating language,

00:39:09.712 Speaker 1 analyzing text, sort of formatting it the way you want it. It's a lot of the magic,

00:39:14.192 Speaker 1 but it's missing some of the, um, precision, maybe better reasoning, a sort

00:39:20.232 Speaker 1 of sequential, like if this, then that logic of a conversation. Um, it has a very

00:39:25.512 Speaker 1 short memory. Um, so that's some of the other limitations to be aware of. One final

00:39:30.352 Speaker 1 limitation to keep in mind is that this thing can only handle 4,000 tokens. So

00:39:36.332 Speaker 1 the GPT-3.5 model is limited in, like, its short-term working memory, how much

00:39:42.292 Speaker 1 it can know about. And the, the amount of space it has to speak back to you

00:39:48.232 Speaker 1 also uses that same shared amount of space. So 4,000 tokens is,

00:39:54.092 Speaker 1 is like 2,500. No, like 3,000 characters usually on average.

00:40:00.092 Speaker 1 So tokens are slightly different than, like, character size. You can, like, use a

00:40:05.772 Speaker 1 calculator to figure out, you know, how they get broken up. I think that's because

00:40:09.732 Speaker 1 some characters end up being condensed, um, down into one, but also white space and

00:40:15.512 Speaker 1 punctuation and stuff like that count. So you have roughly 3,000 characters of

00:40:20.692 Speaker 1 short-term working memory in your prompts. But that amount of space, and you could

00:40:24.812 Speaker 1 see here where I said that this was a learning opportunity, if you try to go above

00:40:28.872 Speaker 1 4,000, it will run out of space to answer you. So whatever you wanna tell this thing

00:40:33.252 Speaker 1 in terms of instructions, it has to fit within that 4,000. Um, so that's the only

00:40:38.552 Speaker 1 other gotcha to keep in mind. So from here, the way that this API works is

00:40:44.612 Speaker 1 you literally just send it this big blob of text, and it sends you back its answer.

00:40:49.252 Speaker 1 Now, in production, you're gonna use stuff like best of, and this is kind of like

00:40:55.992 Speaker 1 this part, right? Where you're trying to pick the best answer to send back. Generate

00:41:00.252 Speaker 1 five different answers and pick the best one. It's going to try to do something like

00:41:04.512 Speaker 1 that, okay? So, but it's more expensive to do that. So this is the general

00:41:10.291 Speaker 1 approach, and hopefully this was enough information that you guys can start

00:41:14.792 Speaker 1 prototyping your own ideas without a developer. Um, and you sort of feel out the

00:41:19.872 Speaker 1 boundaries of how hard is your problem. In this case, scam phone call detection

00:41:24.732 Speaker 1 looks very plausible. Uh, in fact, I would say that these state-of-the-art language

00:41:29.132 Speaker 1 models are exactly what we need to sort of curb this problem, and it's a very

00:41:33.972 Speaker 1 promising startup. And, you know, that is essentially my analysis for them is, yeah,

00:41:38.892 Speaker 1 this looks good. I think we can do this. And the most exciting part is you don't

00:41:42.472 Speaker 1 have to train your own models. This is just an off-the-shelf API from OpenAI. You

00:41:47.532 Speaker 1 can just prompt craft properly and then send it this information, and it acts as a

00:41:53.592 Speaker 1 scam phone call detection tool. Now, you're gonna need to do some extra steps to get

00:41:57.792 Speaker 1 to production to make the software ready. Obviously, there was the phone app and the

00:42:01.252 Speaker 1 other stuff, and there's gonna be some stuff you wanna do to make sure that it's not

00:42:04.832 Speaker 1 giving really bad answers. Maybe that's for another video. But when you're

00:42:08.572 Speaker 1 prototyping something, you don't care about any of that. All you really care about

00:42:12.472 Speaker 1 is, can I get this thing to work well enough that I can prove that it's possible?

00:42:18.472 Speaker 1 You can do that. The rest is basically normal software engineering. It's a solved

00:42:23.132 Speaker 1 problem. So there you have it. I hope that is helpful. Uh, I might prototype some

00:42:28.032 Speaker 1 more things, uh, for ChatGPT. It's not ChatGPT. It's just GPT-3.5. We will get the

00:42:33.572 Speaker 1 ChatGPT API hopefully in a couple weeks, maybe a couple months, and that's a whole

00:42:38.632 Speaker 1 new ballgame 'cause it's even better than this. But you will be far ahead of

00:42:42.732 Speaker 1 understanding how to prototype your ideas and what's possible if you get to work

00:42:46.912 Speaker 1 with just the black part, which, as you can see by the relative sizing, is the most

00:42:51.672 Speaker 1 important part. So don't delay. You can get started now and see what's possible.

00:42:58.832 Speaker 1 Hope that was helpful.

00:43:01.832 Speaker 1 So next time on Seeking Minima, we are going to explore a more

00:43:07.152 Speaker 1 philosophical fun question in the form of a story. So I wanted to try and take a

00:43:12.752 Speaker 1 stab at answering the question: what do we do when we don't have to do

00:43:18.852 Speaker 1 anything? And I tried to do that in a very entertaining way. Let me know what you

00:43:23.652 Speaker 1 think